Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Metagenome-Guided Proteomic Quantification of Reductive Dehalogenases in the Dehalococcoides mccartyi -Containing Consortium SDC-9

At groundwater sites contaminated with chlorinated ethenes, fermentable substrates are often added to promote reductive dehalogenation by indigenous or augmented microorganisms. Contemporary bioremediation performance monitoring relies on nucleic acid biomarkers of key organohalide-respiring bacteria, such as Dehalococcoides mccartyi (Dhc). In this work, metagenome sequencing of the commercial, Dhc-containing consortium, SDC-9, identified 12 reductive dehalogenase (RDase) genes, including pceA (two copies), vcrA, and tceA, and allowed for specific detection and quantification of RDase peptides using liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS). Shotgun (i.e., untargeted) proteomics applied to the SDC-9 consortium grown with tetrachloroethene (PCE) and lactate identified 143 RDase peptides, and 36 distinct peptides that covered greater than 99% of the protein-coding sequences of the PceA, TceA, and VcrA RDases. Quantification of RDase peptides using multiple reaction monitoring (MRM) assays with 13 C-/ 15 N-labeled peptides determined 1.8 × 10 3 TceA and 1.2 × 10 2 VcrA RDase molecules per Dhc cell. Ultimately, the MRM mass spectrometry approach allowed for sensitive detection and accurate quantification of relevant Dhc RDases and has potential utility in bioremediation monitoring regimes.

59 BASIC BIOLOGICAL SCIENCES↗

IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data

Proteoforms, the different forms of a protein with sequence variations including post-translational modifications (PTMs), execute vital functions in biological systems such as cell signaling and epigenetic regulation. Precisely defining the stoichiometry of PTMs has been challenging because, in the widely used bottom-up proteomics methods, the detection occurs at the peptide level and thus the link between peptides and their specific modification site is lost, resulting in proteoform ambiguity. Advances in top-down mass spectrometry (MS) technology have permitted the direct characterization of intact proteoforms and their exact number of modification sites, allowing for the relative quantification of positional isomers (PI). Proteins with positional isomers refers to proteoforms with identical total mass and set of modifications but varying PTM site combinations. The relative abundance of PI can be estimated by matching proteoform-specific fragment ions to top-down tandem MS (MS2) data to localize and quantify modifications. However, current approaches heavily rely on manual annotation. Here, we present IsoForma, an open-source R package for relative quantification of PI within a single tool. We benchmarked IsoForma’s performance against two existing workflows and highlight the similarity of the results and improvements in speed. Overall, IsoForma provides a streamlined process, reduces the time of conducting isoform-based analyses, and offers an essential framework for developing customized proteoform analysis workflows. Finally, the software is open source and available at https://github.com/EMSL-Computing/isoforma-lib.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, overproduction and purification of Vibrio proteolyticus ribosomal protein L18 for in vitro and in vivo studies

A strategy suggested by comparative genomic studies was used to amplify the entire Vibrio proteolyticus (Vp) gene for ribosomal protein L18. Vp L18 and its flanking regions were sequenced and compared with the deduced amino acid (aa) sequences of other known L18 proteins. A 26-aa residue segment at the carboxy terminus contains many strongly conserved residues and may be critical for the L18 interaction with 5S rRNA. This approach should allow rapid characterization of L18 from large numbers of bacteria. Both Vp L18 and Escherichia coli (Ec) L18 were overproduced and purified using a T7 expression vector which fuses an N-terminal peptide segment (His-tag) containing 6 histidine residues to the recombinant protein. The purified fusion proteins, Vp His::L18 and Ec His::L18, were both found to bind to either the Vp 5S or Ec 5S rRNAs in vitro. Vp His::L18 protein was also shown to incorporate into Ec ribosomes in vivo. This His-tag strategy likely will have general applicability for the study of ribosomal proteins in vitro and in vivo.

Non-NASA Center↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology↗

Toho-1 β-lactamase: backbone chemical shift assignments and changes in dynamics upon binding with avibactam

Backbone chemical shift assignments for the Toho-1 β-lactamase (263 amino acids, 28.9 kDa) are reported based on triple resonance solution-state NMR experiments performed on a uniformly 2 H, 13 C, 15 N-labeled sample. These assignments allow for subsequent site-specific characterization at the chemical, structural, and dynamical levels. At the chemical level, titration with the non-β-lactam β-lactamase inhibitor avibactam is found to give chemical shift perturbations indicative of tight covalent binding that allow for mapping of the inhibitor binding site. At the structural level, protein secondary structure is predicted based on the backbone chemical shifts and protein residue sequence using TALOS-N and found to agree well with structural characterization from X-ray crystallography. At the dynamical level, model-free analysis of 15 N relaxation data at a single field of 16.4 T reveals well-ordered structures for the ligand-free and avibactam-bound enzymes with generalized order parameters of ~0.85. Complementary relaxation dispersion experiments indicate that there is an escalation in motions on the millisecond timescale in the vicinity of the active site upon substrate binding. The combination of high rigidity on short timescales and active site flexibility on longer timescales is consistent with hypotheses for achieving both high catalytic efficiency and broad substrate specificity: the induced active site dynamics allows variously sized substrates to be accommodated and increases the probability that the optimal conformation for catalysis will be sampled.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interlaboratory Study for Characterizing Monoclonal Antibodies by Top-Down and Middle-Down Mass Spectrometry

The Consortium for Top-Down Proteomics (www.topdownproteomics.org) launched the present study to assess the state of top-down mass spectrometry (TD MS) and middle-down mass spectrometry (MD MS) for characterizing monoclonal antibody (mAb) primary structures, including their modifications. To meet the needs of the rapidly growing therapeutic antibody market, it is important to develop analytical strategies to characterize the heterogeneity of a therapeutic product’s primary structure reliably and reproducibly. The major objective of the present study was to determine whether current TD/MD MS technologies and protocols can add value to the more commonly employed bottom-up approaches, such as confirming protein integrity, sequencing variable domains, ruling out artefacts, and revealing modifications and their locations. A panel of three mAbs was selected and centrally provided to twenty laboratories worldwide for the analysis: Sigma mAb standard (SiLuLite), NIST mAb standard, and the therapeutic mAb Herceptin (trastuzumab). A variety of different MS instrument platforms and ion dissociation techniques were employed. Overall, the present study confirms that MD/TD MS strategies are valuable and readily available tools in laboratories worldwide. They provide complementary information to the bottom-up approach and can add unique information, which may be crucial for comprehensive mAb characterization. The current limitations, as well as possible solutions to overcome them, are also outlined. One of the limitations revealed by the results of the present study is that the expert knowledge in both experiment and data analysis is indispensable to practice TD/MD MS, but the number of TD/MD MS experts is still very low.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular remodeling in Populus PdKOR RNAi roots profiled using LC-MS/MS proteomics

Plant endo-β-1,4-glucanases belonging to the Glycoside Hydrolase Family 9 have functional roles in cell wall biosynthesis and remodeling via endohydrolysis of (1→4)-β-D-glucosidic linkages. Modification of cell wall chemistry via RNAi-mediated downregulation of Populus deltoides KOR1 (PdKOR), a endo-β-1,4-glucanase gene, in Populus deltoides has been shown to have functional consequences for the composition of secondary metabolome and the ability of modified roots to interact with beneficial microbes. The molecular remodeling that underlies the observed differences at metabolic, physiological, and morphological levels in roots is not well understood. Here we used a LC-MS/MS-based proteome profiling approach to survey the molecular remodeling in root tissues of PdKOR and control plants. A total of 14316 peptides were identified and these mapped to 7139 P. deltoides proteins. Based on 90% sequence identity, the measured protein accessions represent 1187 functional protein groups. Analysis of GO categories and specific individual proteins showed differential expression of proteins relevant to plant-microbe interactions, cell wall chemistry, and metabolism. The new proteome dataset serves as a useful resource for deriving new hypotheses and empirical testing pertaining to functional roles of proteins and pathways in differential priming of plant roots to interactions with microbes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Engineering Pseudomonas putida for production of 3-hydroxyacids using hybrid type I polyketide synthases

Engineered type I polyketide synthases (T1PKSs) are a potentially transformative platform for the biosynthesis of small molecules. Due to their modular nature, T1PKSs can be rationally designed to produce a wide range of bulk or specialty chemicals. While heterologous PKS expression is best studied in microbes of the genus Streptomyces, recent studies have focused on the exploration of non-native PKS hosts. The biotechnological production of chemicals in fast growing and industrial relevant hosts has numerous economic and logistic advantages. With its native ability to utilize alternative feedstocks, Pseudomonas putida has emerged as a promising workhorse for the sustainable production of small molecules. Here, we outline the assessment of P. putida as a host for the expression of engineered T1PKSs and production of 3-hydroxyacids. After establishing the functional expression of an engineered T1PKS, we successfully expanded and increased the pool of available acyl-CoAs needed for the synthesis of polyketides using transposon sequencing and protein degradation tagging. This work demonstrates the potential of T1PKSs in P. putida as a production platform for the sustainable biosynthesis of unnatural polyketides.

Schmidt, Matthias↗

Computational Prediction of Coiled–Coil Protein Gelation Dynamics and Structure

Protein hydrogels represent an important and growing biomaterial for a multitude of applications, including diagnostics and drug delivery. We have previously explored the ability to engineer the thermoresponsive supramolecular assembly of coiled–coil proteins into hydrogels with varying gelation properties, where we have defined important parameters in the coiled–coil hydrogel design. Using Rosetta energy scores and Poisson–Boltzmann electrostatic energies, we iterate a computational design strategy to predict the gelation of coiled–coil proteins while simultaneously exploring five new coiled–coil protein hydrogel sequences. Provided this library, we explore the impact of in silico energies on structure and gelation kinetics, where we also reveal a range of blue autofluorescence that enables hydrogel disassembly and recovery. As a result of this library, we identify the new coiled–coil hydrogel sequence, Q5, capable of gelation within 24 h at 4 °C, a more than 2-fold increase over that of our previous iteration Q2. The fast gelation time of Q5 enables the assessment of structural transition in real time using small-angle X-ray scattering (SAXS) that is correlated to coarse-grained and atomistic molecular dynamics simulations revealing the supramolecular assembling behavior of coiled–coils toward nanofiber assembly and gelation. This work represents the first system of hydrogels with predictable self-assembly, autofluorescent capability, and a molecular model of coiled–coil fiber formation.

36 MATERIALS SCIENCE↗

Directing Nanoparticle Organization in Response to Diverse Chemical Inputs

Signaling cascades are crucial for transducing stimuli in biological systems, enabling multiple stimuli to regulate a downstream target with precisely controlled timing and amplifying signals through a series of intermediary reactions. Developing a robust signaling system with such capabilities would be pivotal for programming complex behaviors in synthetic DNA-based molecular devices. However, although “software” such as nucleic acid circuits could potentially be harnessed to relay signals to DNA-based nanostructure hardware, such explorations have been limited. Here, in this study, we develop a platform for transducing a variety of stimuli via messenger-mediated reactions to regulate the release and reloading of gold nanoparticles (AuNPs) in a 3D DNA framework. In the first step, an in vitro transcription circuit is engineered to sense and amplify chemical stimuli, including arbitrary DNA sequences and proteins, producing RNA. In the second step, the RNA releases the DNA-coated AuNPs from the DNA framework via a strand displacement reaction. AuNP reloading is controlled by a separate step driven by degradation of the RNA. Our platform holds promise for applications requiring dynamic multiagent control over DNA-based devices, offering a versatile tool for advanced molecular device engineering.

36 MATERIALS SCIENCE↗

Viruses infecting a warm water picoeukaryote shed light on spatial co-occurrence dynamics of marine viruses and their hosts

The marine picoeukaryote Bathycoccus prasinos has been considered a cosmopolitan alga, although recent studies indicate two ecotypes exist, Clade BI (B. prasinos) and Clade BII. Viruses that infect Bathycoccus Clade BI are known (BpVs), but not that infect BII. We isolated three dsDNA prasinoviruses from the Sargasso Sea against Clade BII isolate RCC716. The BII-Vs do not infect BI, and two (BII-V2 and BII-V3) have larger genomes (~210 kb) than BI-Viruses and BII-V1. BII-Vs share ~90% of their proteins, and between 65% to 83% of their proteins with sequenced BpVs. Phylogenomic reconstructions and PolB analyses establish close-relatedness of BII-V2 and BII-V3, yet BII-V2 has 10-fold higher infectivity and induces greater mortality on host isolate RCC716. BII-V1 is more distant, has a shorter latent period, and infects both available BII isolates, RCC716 and RCC715, while BII-V2 and BII-V3 do not exhibit productive infection of the latter in our experiments. Global metagenome analyses show Clade BI and BII algal relative abundances correlate positively with their respective viruses. The distributions delineate BI/BpVs as occupying lower temperature mesotrophic and coastal systems, whereas BII/BII-Vs occupy warmer temperature, higher salinity ecosystems. Accordingly, with molecular diagnostic support, we name Clade BII Bathycoccus calidus sp. nov. and propose that molecular diversity within this new species likely connects to the differentiated host-virus dynamics observed in our time course experiments. Overall, the tightly linked biogeography of Bathycoccus host and virus clades observed herein supports species-level host specificity, with strain-level variations in infection parameters.

59 BASIC BIOLOGICAL SCIENCES↗

Diploid genomic architecture of Nitzschia inconspicua, an elite biomass production diatom

Abstract A near-complete diploid nuclear genome and accompanying circular mitochondrial and chloroplast genomes have been assembled from the elite commercial diatom species Nitzschia inconspicua . The 50 Mbp haploid size of the nuclear genome is nearly double that of model diatom Phaeodactylum tricornutum , but 30% smaller than closer relative Fragilariopsis cylindrus . Diploid assembly, which was facilitated by low levels of allelic heterozygosity (2.7%), included 14 candidate chromosome pairs composed of long, syntenic contigs, covering 93% of the total assembly. Telomeric ends were capped with an unusual 12-mer, G-rich, degenerate repeat sequence. Predicted proteins were highly enriched in strain-specific marker domains associated with cell-surface adhesion, biofilm formation, and raphe system gliding motility. Expanded species-specific families of carbonic anhydrases suggest potential enhancement of carbon concentration efficiency, and duplicated glycolysis and fatty acid synthesis pathways across cytosolic and organellar compartments may enhance peak metabolic output, contributing to competitive success over other organisms in mixed cultures. The N. inconspicua genome delivers a robust new reference for future functional and transcriptomic studies to illuminate the physiology of benthic pennate diatoms and harness their unique adaptations to support commercial algae biomass and bioproduct production.

09 BIOMASS FUELS↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Novel symmetry-preserving neural network model for phylogenetic inference

Abstract Motivation Scientists world-wide are putting together massive efforts to understand how the biodiversity that we see on Earth evolved from single-cell organisms at the origin of life and this diversification process is represented through the Tree of Life. Low sampling rates and high heterogeneity in the rate of evolution across sites and lineages produce a phenomenon denoted “long branch attraction” (LBA) in which long nonsister lineages are estimated to be sisters regardless of their true evolutionary relationship. LBA has been a pervasive problem in phylogenetic inference affecting different types of methodologies from distance-based to likelihood-based. Results Here, we present a novel neural network model that outperforms standard phylogenetic methods and other neural network implementations under LBA settings. Furthermore, unlike existing neural network models in phylogenetics, our model naturally accounts for the tree isomorphisms via permutation invariant functions which ultimately result in lower memory and allows the seamless extension to larger trees. Availability and implementation We implement our novel theory on an open-source publicly available GitHub repository: https://github.com/crsl4/nn-phylogenetics.

59 BASIC BIOLOGICAL SCIENCES↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

Persistent minimal sequences of SARS-CoV-2

Abstract Motivation Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has caused more than 14 million cases and more than half million deaths. Given the absence of implemented therapies, new analysis, diagnosis and therapeutics are of great importance. Results Analysis of SARS-CoV-2 genomes from the current outbreak reveals the presence of short persistent DNA/RNA sequences that are absent from the human genome and transcriptome (PmRAWs). For the PmRAWs with length 12, only four exist at the same location in all SARS-CoV-2. At the gene level, we found one PmRAW of size 13 at the Spike glycoprotein coding sequence. This protein is fundamental for binding in human ACE2 and further use as an entry receptor to invade target cells. Applying protein structural prediction, we localized this PmRAW at the surface of the Spike protein, providing a potential targeted vector for diagnostics and therapeutics. In addition, we show a new pattern of relative absent words (RAWs), characterized by the progressive increase of GC content (Guanine and Cytosine) according to the decrease of RAWs length, contrarily to the virus and host genome distributions. New analysis shows the same property during the Ebola virus outbreak. At a computational level, we improved the alignment-free method to identify pathogen-specific signatures in balance with GC measures and removed previous size limitations. Availability and implementation https://github.com/cobilab/eagle. Supplementary information Supplementary data are available at Bioinformatics online.

Pratas, Diogo↗