Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Toho-1 β-lactamase: backbone chemical shift assignments and changes in dynamics upon binding with avibactam

Backbone chemical shift assignments for the Toho-1 β-lactamase (263 amino acids, 28.9 kDa) are reported based on triple resonance solution-state NMR experiments performed on a uniformly 2 H, 13 C, 15 N-labeled sample. These assignments allow for subsequent site-specific characterization at the chemical, structural, and dynamical levels. At the chemical level, titration with the non-β-lactam β-lactamase inhibitor avibactam is found to give chemical shift perturbations indicative of tight covalent binding that allow for mapping of the inhibitor binding site. At the structural level, protein secondary structure is predicted based on the backbone chemical shifts and protein residue sequence using TALOS-N and found to agree well with structural characterization from X-ray crystallography. At the dynamical level, model-free analysis of 15 N relaxation data at a single field of 16.4 T reveals well-ordered structures for the ligand-free and avibactam-bound enzymes with generalized order parameters of ~0.85. Complementary relaxation dispersion experiments indicate that there is an escalation in motions on the millisecond timescale in the vicinity of the active site upon substrate binding. The combination of high rigidity on short timescales and active site flexibility on longer timescales is consistent with hypotheses for achieving both high catalytic efficiency and broad substrate specificity: the induced active site dynamics allows variously sized substrates to be accommodated and increases the probability that the optimal conformation for catalysis will be sampled.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interlaboratory Study for Characterizing Monoclonal Antibodies by Top-Down and Middle-Down Mass Spectrometry

The Consortium for Top-Down Proteomics (www.topdownproteomics.org) launched the present study to assess the state of top-down mass spectrometry (TD MS) and middle-down mass spectrometry (MD MS) for characterizing monoclonal antibody (mAb) primary structures, including their modifications. To meet the needs of the rapidly growing therapeutic antibody market, it is important to develop analytical strategies to characterize the heterogeneity of a therapeutic product’s primary structure reliably and reproducibly. The major objective of the present study was to determine whether current TD/MD MS technologies and protocols can add value to the more commonly employed bottom-up approaches, such as confirming protein integrity, sequencing variable domains, ruling out artefacts, and revealing modifications and their locations. A panel of three mAbs was selected and centrally provided to twenty laboratories worldwide for the analysis: Sigma mAb standard (SiLuLite), NIST mAb standard, and the therapeutic mAb Herceptin (trastuzumab). A variety of different MS instrument platforms and ion dissociation techniques were employed. Overall, the present study confirms that MD/TD MS strategies are valuable and readily available tools in laboratories worldwide. They provide complementary information to the bottom-up approach and can add unique information, which may be crucial for comprehensive mAb characterization. The current limitations, as well as possible solutions to overcome them, are also outlined. One of the limitations revealed by the results of the present study is that the expert knowledge in both experiment and data analysis is indispensable to practice TD/MD MS, but the number of TD/MD MS experts is still very low.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular remodeling in Populus PdKOR RNAi roots profiled using LC-MS/MS proteomics

Plant endo-β-1,4-glucanases belonging to the Glycoside Hydrolase Family 9 have functional roles in cell wall biosynthesis and remodeling via endohydrolysis of (1→4)-β-D-glucosidic linkages. Modification of cell wall chemistry via RNAi-mediated downregulation of Populus deltoides KOR1 (PdKOR), a endo-β-1,4-glucanase gene, in Populus deltoides has been shown to have functional consequences for the composition of secondary metabolome and the ability of modified roots to interact with beneficial microbes. The molecular remodeling that underlies the observed differences at metabolic, physiological, and morphological levels in roots is not well understood. Here we used a LC-MS/MS-based proteome profiling approach to survey the molecular remodeling in root tissues of PdKOR and control plants. A total of 14316 peptides were identified and these mapped to 7139 P. deltoides proteins. Based on 90% sequence identity, the measured protein accessions represent 1187 functional protein groups. Analysis of GO categories and specific individual proteins showed differential expression of proteins relevant to plant-microbe interactions, cell wall chemistry, and metabolism. The new proteome dataset serves as a useful resource for deriving new hypotheses and empirical testing pertaining to functional roles of proteins and pathways in differential priming of plant roots to interactions with microbes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Engineering Pseudomonas putida for production of 3-hydroxyacids using hybrid type I polyketide synthases

Engineered type I polyketide synthases (T1PKSs) are a potentially transformative platform for the biosynthesis of small molecules. Due to their modular nature, T1PKSs can be rationally designed to produce a wide range of bulk or specialty chemicals. While heterologous PKS expression is best studied in microbes of the genus Streptomyces, recent studies have focused on the exploration of non-native PKS hosts. The biotechnological production of chemicals in fast growing and industrial relevant hosts has numerous economic and logistic advantages. With its native ability to utilize alternative feedstocks, Pseudomonas putida has emerged as a promising workhorse for the sustainable production of small molecules. Here, we outline the assessment of P. putida as a host for the expression of engineered T1PKSs and production of 3-hydroxyacids. After establishing the functional expression of an engineered T1PKS, we successfully expanded and increased the pool of available acyl-CoAs needed for the synthesis of polyketides using transposon sequencing and protein degradation tagging. This work demonstrates the potential of T1PKSs in P. putida as a production platform for the sustainable biosynthesis of unnatural polyketides.

Schmidt, Matthias↗

Computational Prediction of Coiled–Coil Protein Gelation Dynamics and Structure

Protein hydrogels represent an important and growing biomaterial for a multitude of applications, including diagnostics and drug delivery. We have previously explored the ability to engineer the thermoresponsive supramolecular assembly of coiled–coil proteins into hydrogels with varying gelation properties, where we have defined important parameters in the coiled–coil hydrogel design. Using Rosetta energy scores and Poisson–Boltzmann electrostatic energies, we iterate a computational design strategy to predict the gelation of coiled–coil proteins while simultaneously exploring five new coiled–coil protein hydrogel sequences. Provided this library, we explore the impact of in silico energies on structure and gelation kinetics, where we also reveal a range of blue autofluorescence that enables hydrogel disassembly and recovery. As a result of this library, we identify the new coiled–coil hydrogel sequence, Q5, capable of gelation within 24 h at 4 °C, a more than 2-fold increase over that of our previous iteration Q2. The fast gelation time of Q5 enables the assessment of structural transition in real time using small-angle X-ray scattering (SAXS) that is correlated to coarse-grained and atomistic molecular dynamics simulations revealing the supramolecular assembling behavior of coiled–coils toward nanofiber assembly and gelation. This work represents the first system of hydrogels with predictable self-assembly, autofluorescent capability, and a molecular model of coiled–coil fiber formation.

36 MATERIALS SCIENCE↗

Directing Nanoparticle Organization in Response to Diverse Chemical Inputs

Signaling cascades are crucial for transducing stimuli in biological systems, enabling multiple stimuli to regulate a downstream target with precisely controlled timing and amplifying signals through a series of intermediary reactions. Developing a robust signaling system with such capabilities would be pivotal for programming complex behaviors in synthetic DNA-based molecular devices. However, although “software” such as nucleic acid circuits could potentially be harnessed to relay signals to DNA-based nanostructure hardware, such explorations have been limited. Here, in this study, we develop a platform for transducing a variety of stimuli via messenger-mediated reactions to regulate the release and reloading of gold nanoparticles (AuNPs) in a 3D DNA framework. In the first step, an in vitro transcription circuit is engineered to sense and amplify chemical stimuli, including arbitrary DNA sequences and proteins, producing RNA. In the second step, the RNA releases the DNA-coated AuNPs from the DNA framework via a strand displacement reaction. AuNP reloading is controlled by a separate step driven by degradation of the RNA. Our platform holds promise for applications requiring dynamic multiagent control over DNA-based devices, offering a versatile tool for advanced molecular device engineering.

36 MATERIALS SCIENCE↗

Viruses infecting a warm water picoeukaryote shed light on spatial co-occurrence dynamics of marine viruses and their hosts

The marine picoeukaryote Bathycoccus prasinos has been considered a cosmopolitan alga, although recent studies indicate two ecotypes exist, Clade BI (B. prasinos) and Clade BII. Viruses that infect Bathycoccus Clade BI are known (BpVs), but not that infect BII. We isolated three dsDNA prasinoviruses from the Sargasso Sea against Clade BII isolate RCC716. The BII-Vs do not infect BI, and two (BII-V2 and BII-V3) have larger genomes (~210 kb) than BI-Viruses and BII-V1. BII-Vs share ~90% of their proteins, and between 65% to 83% of their proteins with sequenced BpVs. Phylogenomic reconstructions and PolB analyses establish close-relatedness of BII-V2 and BII-V3, yet BII-V2 has 10-fold higher infectivity and induces greater mortality on host isolate RCC716. BII-V1 is more distant, has a shorter latent period, and infects both available BII isolates, RCC716 and RCC715, while BII-V2 and BII-V3 do not exhibit productive infection of the latter in our experiments. Global metagenome analyses show Clade BI and BII algal relative abundances correlate positively with their respective viruses. The distributions delineate BI/BpVs as occupying lower temperature mesotrophic and coastal systems, whereas BII/BII-Vs occupy warmer temperature, higher salinity ecosystems. Accordingly, with molecular diagnostic support, we name Clade BII Bathycoccus calidus sp. nov. and propose that molecular diversity within this new species likely connects to the differentiated host-virus dynamics observed in our time course experiments. Overall, the tightly linked biogeography of Bathycoccus host and virus clades observed herein supports species-level host specificity, with strain-level variations in infection parameters.

59 BASIC BIOLOGICAL SCIENCES↗

Diploid genomic architecture of Nitzschia inconspicua, an elite biomass production diatom

Abstract A near-complete diploid nuclear genome and accompanying circular mitochondrial and chloroplast genomes have been assembled from the elite commercial diatom species Nitzschia inconspicua . The 50 Mbp haploid size of the nuclear genome is nearly double that of model diatom Phaeodactylum tricornutum , but 30% smaller than closer relative Fragilariopsis cylindrus . Diploid assembly, which was facilitated by low levels of allelic heterozygosity (2.7%), included 14 candidate chromosome pairs composed of long, syntenic contigs, covering 93% of the total assembly. Telomeric ends were capped with an unusual 12-mer, G-rich, degenerate repeat sequence. Predicted proteins were highly enriched in strain-specific marker domains associated with cell-surface adhesion, biofilm formation, and raphe system gliding motility. Expanded species-specific families of carbonic anhydrases suggest potential enhancement of carbon concentration efficiency, and duplicated glycolysis and fatty acid synthesis pathways across cytosolic and organellar compartments may enhance peak metabolic output, contributing to competitive success over other organisms in mixed cultures. The N. inconspicua genome delivers a robust new reference for future functional and transcriptomic studies to illuminate the physiology of benthic pennate diatoms and harness their unique adaptations to support commercial algae biomass and bioproduct production.

09 BIOMASS FUELS↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Novel symmetry-preserving neural network model for phylogenetic inference

Abstract Motivation Scientists world-wide are putting together massive efforts to understand how the biodiversity that we see on Earth evolved from single-cell organisms at the origin of life and this diversification process is represented through the Tree of Life. Low sampling rates and high heterogeneity in the rate of evolution across sites and lineages produce a phenomenon denoted “long branch attraction” (LBA) in which long nonsister lineages are estimated to be sisters regardless of their true evolutionary relationship. LBA has been a pervasive problem in phylogenetic inference affecting different types of methodologies from distance-based to likelihood-based. Results Here, we present a novel neural network model that outperforms standard phylogenetic methods and other neural network implementations under LBA settings. Furthermore, unlike existing neural network models in phylogenetics, our model naturally accounts for the tree isomorphisms via permutation invariant functions which ultimately result in lower memory and allows the seamless extension to larger trees. Availability and implementation We implement our novel theory on an open-source publicly available GitHub repository: https://github.com/crsl4/nn-phylogenetics.

59 BASIC BIOLOGICAL SCIENCES↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

Persistent minimal sequences of SARS-CoV-2

Abstract Motivation Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has caused more than 14 million cases and more than half million deaths. Given the absence of implemented therapies, new analysis, diagnosis and therapeutics are of great importance. Results Analysis of SARS-CoV-2 genomes from the current outbreak reveals the presence of short persistent DNA/RNA sequences that are absent from the human genome and transcriptome (PmRAWs). For the PmRAWs with length 12, only four exist at the same location in all SARS-CoV-2. At the gene level, we found one PmRAW of size 13 at the Spike glycoprotein coding sequence. This protein is fundamental for binding in human ACE2 and further use as an entry receptor to invade target cells. Applying protein structural prediction, we localized this PmRAW at the surface of the Spike protein, providing a potential targeted vector for diagnostics and therapeutics. In addition, we show a new pattern of relative absent words (RAWs), characterized by the progressive increase of GC content (Guanine and Cytosine) according to the decrease of RAWs length, contrarily to the virus and host genome distributions. New analysis shows the same property during the Ebola virus outbreak. At a computational level, we improved the alignment-free method to identify pathogen-specific signatures in balance with GC measures and removed previous size limitations. Availability and implementation https://github.com/cobilab/eagle. Supplementary information Supplementary data are available at Bioinformatics online.

Pratas, Diogo↗

Codon2Vec v1.0

Background: Codon2Vec is an embedding neural network that predicts 'high' or 'low' gene expression directly from the protein-coding sequences. Embedding neural networks are commonly used for natural language processing (NLP) applications. Analogous to how an English sentence is a string of words, a gene can be thought of as a string of codons. Similar to how NLP neural networks model English sentences as a non-random sequence of words, we considered a coding sequence as a non-random non-overlapping array of codons (k-mers of length = 3). Value Proposition: - Codon2Vec achieved a high median AUC-ROC score of 83.8% when trained and applied to transcriptomic data from 300 fungal species - Unlike Codo2Vec, conventional methods predicting for expression based on codon usage rely on a priori knowledge of optimal codons or a set of reference genes. - Unlike Codon2vec, these methods do not account for the effect of codon order on gene expression. - Codon2Vec neural network bypasses the need for artisanal feature selection step that is necessary for traditional machine learning models.

Wint, Rhondene↗

Functional Diversification of Populus FLOWERING LOCUS D-LIKE3 Transcription Factor and Two Paralogs in Shoot Ontogeny, Flowering, and Vegetative Phenology

Both the evolution of tree taxa and whole-genome duplication (WGD) have occurred many times during angiosperm evolution. Transcription factors are preferentially retained following WGD suggesting that functional divergence of duplicates could contribute to traits distinctive to the tree growth habit. We used gain- and loss-of-function transgenics, photoperiod treatments, and circannual expression studies in adult trees to study the diversification of three Populus FLOWERING LOCUS D-LIKE (FDL) genes encoding bZIP transcription factors. Expression patterns and transgenic studies indicate that FDL2.2 promotes flowering and that FDL1 and FDL3 function in different vegetative phenophases. Study of dominant repressor FDL versions indicates that the FDL proteins are partially equivalent in their ability to alter shoot growth. Like its paralogs, FDL3 overexpression delays short day-induced growth cessation, but also induces distinct heterochronic shifts in shoot development—more rapid phytomer initiation and coordinated delay in both leaf expansion and the transition to secondary growth in long days, but not in short days. Our results indicate that both regulatory and protein coding sequence variation contributed to diversification of FDL paralogs that has led to a degree of specialization in multiple developmental processes important for trees and their local adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

Differential transcriptomic alterations in nasal versus lung tissue of acrolein-exposed rats

Introduction: Acrolein is a significant component of anthropogenic and wildfire emissions, as well as cigarette smoke. Although acrolein primarily deposits in the upper respiratory tract upon inhalation, patterns of site-specific injury in nasal versus pulmonary tissues are not well characterized. This assessment is critical in the design of in vitro and in vivo studies performed for assessing health risk of irritant air pollutants. Methods: In this study, male and female Wistar-Kyoto rats were exposed nose-only to air or acrolein. Rats in the acrolein exposure group were exposed to incremental concentrations of acrolein (0, 0.1, 0.316, 1 ppm) for the first 30 min, followed by a 3.5 h exposure at 3.16 ppm. In the first cohort of male and female rats, nasal and bronchoalveolar lavage fluids were analyzed for markers of inflammation, and in a second cohort of males, nasal airway and left lung tissues were used for mRNA sequencing. Results: Protein leakage in nasal airways of acrolein-exposed rats was similar in both sexes; however, inflammatory cells and cytokine increases were more pronounced in males when compared to females. No consistent changes were noted in bronchoalveolar lavage fluid of males or females except for increases in total cells and IL-6. Acrolein-exposed male rats had 452 differentially expressed genes (DEGs) in nasal tissue versus only 95 in the lung. Pathway analysis of DEGs in the nose indicated acute phase response signaling, Nrf2-mediated oxidative stress, unfolded protein response, and other inflammatory pathways, whereas in the lung, xenobiotic metabolism pathways were changed. Genes associated with glucocorticoid and GPCR signaling were also changed in the nose but not in the lung. Discussion: These data provide insights into inhaled acrolein-mediated sex-specific injury/inflammation in the nasal and pulmonary airways. The transcriptional response in the nose reflects acrolein-induced acute oxidative and cytokine signaling changes, which might have implications for upper airway inflammatory disease susceptibility.

Alewel, Devin I.↗

TopPICR: A Companion R Package for Top-Down Proteomics Data Analysis

Top-down proteomics is the analysis of proteins in their intact form without proteolysis, thus preserving valuable information about post-translational modifications, isoforms, and proteolytic processing. However, it is still a developing field due to limitations in the instrumentation, difficulties with interpretation of complex mass spectra, and a lack of well-established quantification approaches. TopPIC is one of the popular tools for proteoform identification. Here we extended its capabilities into label-free proteoform quantification by developing a companion R package (TopPICR). Key steps in the TopPICR pipeline include filtering identifications, inferring a minimal set of protein accessions explaining the observed sequences, aligning retention times, recalibrating measured masses, clustering features across datasets, and finally compiling feature intensities using the match-between-runs approach. The output of the pipeline is an MSnSet object which makes downstream data analysis seamlessly compatible with packages from the Bioconductor project. It also provides the capability for visualizing proteoforms within the context of the parent protein sequence. The functionality of TopPICR is demonstrated on top-down LC-MS/MS datasets of 10 human-in-mouse xenografts of luminal and basal breast tumor samples.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗