Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Proteome-scale Deployment of Protein Structure Prediction Workflows on the Summit Supercomputer

Deep learning has contributed to major advances in the prediction of protein structure from sequence, a fundamental problem in structural bioinformatics. With predictions now approaching the accuracy of crystallographic experiments, and with accelerators like GPUs and TPUs making inference using large models rapid, genome-level structure prediction becomes an obvious aim. Leadership-class computing resources can be used to perform genome-scale protein structure prediction using state-of-the-art deep learning models, providing a wealth of new data for systems biology applications. Here we describe our efforts to efficiently deploy the AlphaFold v.2 program, for full-proteome structure prediction, at scale on the Oak Ridge Leadership Computing Facility's resources, including the Summit supercomputer. We performed inference to produce the predicted structures for 40,526 protein sequences, corresponding to four prokaryotic proteomes and one plant proteome, using under 4,400 total Summit node hours, equivalent to using the majority of the supercomputer for a little over one hour. We also designed an optimized structure refinement that reduced the time for the relaxation stage of the AlphaFold pipeline by over 10X for longer sequences. We demonstrate the types of analyses that can be performed on proteome-scale collections of sequences, including a search for novel quaternary structures and implications for functional annotation.

Gao, Mu↗

luoyunan/ECNet: First release

ECNet (evolutionary context-integrated neural network) is a deep learning model that guides protein engineering by predicting protein fitness from the sequence. It integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe.

Luo, Yunan↗

SCOPe: improvements to the structural classification of proteins – extended database to facilitate variant interpretation and machine learning

Abstract The Structural Classification of Proteins—extended (SCOPe, https://scop.berkeley.edu) knowledgebase aims to provide an accurate, detailed, and comprehensive description of the structural and evolutionary relationships amongst the majority of proteins of known structure, along with resources for analyzing the protein structures and their sequences. Structures from the PDB are divided into domains and classified using a combination of manual curation and highly precise automated methods. In the current release of SCOPe, 2.08, we have developed search and display tools for analysis of genetic variants we mapped to structures classified in SCOPe. In order to improve the utility of SCOPe to automated methods such as deep learning classifiers that rely on multiple alignment of sequences of homologous proteins, we have introduced new machine-parseable annotations that indicate aberrant structures as well as domains that are distinguished by a smaller repeat unit. We also classified structures from 74 of the largest Pfam families not previously classified in SCOPe, and we improved our algorithm to remove N- and C-terminal cloning, expression and purification sequences from SCOPe domains. SCOPe 2.08-stable classifies 106 976 PDB entries (about 60% of PDB entries).

59 BASIC BIOLOGICAL SCIENCES↗

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Hemiptera phylogenomic resources: Tree‐based orthology prediction and conserved exon identification

Abstract High‐throughput sequencing of transcriptomes and targeted genomic regions are advancing our knowledge of The Tree of Life. Building phylogenies with regions of the genome requires 1‐to‐1 orthologue resources of genes and noncoding loci. One organismal group that has received little attention in this area is the Hemiptera, the fifth largest insect order represented by ~103,590 named species. Here, we present a set of 3,872 Hemiptera 1‐to‐1 orthogroups based on tree‐based orthology inference of eight Hemiptera species with publicly available genome sequences. We also estimate a set of 406 orthologous exons with similar mRNA splice sites that can be used for Sanger sequencing and develop enrichment probes for targeted genome sequencing for phylogenomic inference. We show this novel set of orthologues is informative at the protein, coding sequence and exon molecular levels and provides robust branch support in both gene tree–species tree methods and concatenated sequence phylogenies. In addition, we demonstrate the utility of these loci to resolve relationships in whiteflies, Bemisia tabaci , a large species complex with few phylogenomic resources. Last, we compare our Hemiptera phylogeny with previously published phylogenies and other orthologue databases, while providing suggestions on further improvement to this phylogenomic resource.

Owen, Christopher L.↗

Top-down mass spectrometry and assigning internal fragments for determining disulfide bond positions in proteins

Disulfide bonds in proteins have a substantial impact on protein structure, stability, and biological activity. Localizing disulfide bonds is critical for understanding protein folding and higher-order structure. Conventional top-down mass spectrometry (TD-MS), where only terminal fragments are assigned for disulfide-intact proteins, can access disulfide information, but suffers from low fragmentation efficiency, thereby limiting sequence coverage. Here, we show that assigning internal fragments generated from TD-MS enhances the sequence coverage of disulfide-intact proteins by 20–60% by returning information from the interior of the protein sequence, which cannot be obtained by terminal fragments alone. Further, the inclusion of internal fragments can extend the sequence information of disulfide-intact proteins to near complete sequence coverage. Importantly, the enhanced sequence information that arise from the assignment of internal fragments can be used to determine the relative position of disulfide bonds and the exact disulfide connectivity between cysteines. The data presented here demonstrates the benefits of incorporating internal fragment analysis into the TD-MS workflow for analyzing disulfide-intact proteins, which would be valuable for characterizing biotherapeutic proteins such as monoclonal antibodies and antibody–drug conjugates.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Complete Genome Sequence of Acidithiobacillus ferridurans JAGS, Isolated from Acidic Mine Drainage

We report a complete genome sequence of Acidithiobacillus ferridurans JAGS, determined using PacBio single-molecule real-time (SMRT) sequencing. The circular genome of JAGS (2,933,811 bp; GC content, 58.57%) contains 3,001 protein-coding sequences, 46 tRNAs, and 6 rRNAs. Predicted genes indicate the potential to fix CO 2 and N 2 and to utilize Fe 2+ , S0, and H 2 as energy sources.

59 BASIC BIOLOGICAL SCIENCES↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

Complete Genome Sequence of Sulfurospirillum sp. Strain ACS DCE , an Anaerobic Bacterium That Respires Tetrachloroethene under Acidic pH Conditions

Sulfurospirillum sp. strain ACS DCE couples growth with reductive dechlorination of tetrachloroethene to cis-1,2-dichloroethene at pH values as low as 5.5. The genome sequence of strain ACS DCE consists of a circular 2,737,849-bp chromosome and a 39,868-bp plasmid and carries 2,737 protein-coding sequences, including two reductive dehalogenase genes.

59 BASIC BIOLOGICAL SCIENCES↗

Towards predictive control of reversible nanoparticle assembly with solid-binding proteins

Although a broad range of ligand-functionalized nanoparticles and physico-chemical triggers have been exploited to create stimuli-responsive colloidal systems, little attention has been paid to the reversible assembly of unmodified nanoparticles with non-covalently bound proteins. Previously, we reported that a derivative of green fluorescent protein engineered with oppositely located silica-binding peptides mediates the repeated assembly and disassembly of 10-nm silica nanoparticles when pH is toggled between 7.5 and 8.5. We captured the subtle interplay between interparticle electrostatic repulsion and their protein-mediated short-range attraction with a multiscale model energetically benchmarked to collective system behavior captured by scattering experiments. Here, in this work, we show that both solution conditions (pH and ionic strength) and protein engineering (sequence and position of engineered silica-binding peptides) provide pathways for reversible control over growth and fragmentation, leading to clusters ranging in size from 25 nm protein-coated particles to micrometer-size aggregate. We further find that the higher electrolyte environment associated with successive cycles of base addition eventually eliminates reversibility. Our model accurately predicts these multiple length scales phenomena. The underpinning concepts provide design principles for the dynamic control of other protein- and particle-based nanocomposites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Suppressor Mutations in Type II Secretion Mutants of Vibrio cholerae : Inactivation of the VesC Protease

The type II secretion system (T2SS) is a conserved transport pathway responsible for the secretion of a range of virulence factors by many pathogens, including Vibrio cholerae. Disruption of the T2SS genes in V. cholerae results in loss of secretion, changes in cell envelope function, and growth defects. While T2SS mutants are viable, high-throughput genomic analyses have listed these genes among essential genes. To investigate whether secondary mutations arise as a consequence of T2SS inactivation, we sequenced the genomes of six V. cholerae T2SS mutants with deletions or insertions in either the epsG, epsL, or epsM genes and identified secondary mutations in all mutants. Two of the six T2SS mutants contain distinct mutations in the gene encoding the T2SS-secreted protease VesC. Other mutations were found in genes coding for V. cholerae cell envelope proteins. Subsequent sequence analysis of the vesC gene in 92 additional T2SS mutant isolates identified another 19 unique mutations including insertions or deletions, sequence duplications, and single-nucleotide changes resulting in amino acid substitutions in the VesC protein. Analysis of VesC variants and the X-ray crystallographic structure of wild-type VesC suggested that all mutations lead to loss of VesC production and/or function. One possible mechanism by which V. cholerae T2SS mutagenesis can be tolerated is through selection of vesC-inactivating mutations, which may, in part, suppress cell envelope damage, establishing permissive conditions for the disruption of the T2SS. Other mutations may have been acquired in genes encoding essential cell envelope proteins to prevent proteolysis by VesC.

59 BASIC BIOLOGICAL SCIENCES↗

Expression of blue pigment synthetase a from Streptomyces lavenduale reveals insights on the effects of refactoring biosynthetic megasynthases for heterologous expression in Escherichia coli .

High GC bacteria from the genus Streptomyces harbor expansive secondary metabolism. The expression of biosynthetic proteins and the characterization and identification of biological "parts" for synthetic biology purposes from such pathways are of interest. However, the high GC content of proteins from actinomycetes in addition to the large size and multi-domain architecture of many biosynthetic proteins (such as non-ribosomal peptide synthetases; NRPSs, and polyketide synthases; PKSs often called "megasynthases") often presents issues with full-length translation and folding. Here we evaluate a non-ribosomal peptide synthetase (NRPS) from Streptomyces lavenduale, a multidomain "megasynthase" gene that comes from a high GC (72.5%) genome. While a preliminary step in revealing differences, to our knowledge this presents the first head-to-head comparison of codon-optimized sequences versus a native sequence of proteins of streptomycete origin heterologously expressed in E. coli. We found that any disruption in co-translational folding from codon mismatch that reduces the titer of indigoidine is explainable via the formation of more inclusion bodies as opposed to compromising folding or posttranslational modification in the soluble fraction. In conclusion, this result supports that one could apply any refactoring strategies that improve soluble expression in E. coli without concern that the protein that reaches the soluble fraction is differentially folded.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence specificity analysis of the SETD2 protein lysine methyltransferase and discovery of a SETD2 super-substrate

SETD2 catalyzes methylation at lysine 36 of histone H3 and it has many disease connections. We investigated the substrate sequence specificity of SETD2 and identified nine additional peptide and one protein (FBN1) substrates. Our data showed that SETD2 strongly prefers amino acids different from those in the H3K36 sequence at several positions of its specificity profile. Based on this, we designed an optimized super-substrate containing four amino acid exchanges and show by quantitative methylation assays with SETD2 that the super-substrate peptide is methylated about 290-fold more efficiently than the H3K36 peptide. Protein methylation studies confirmed very strong SETD2 methylation of the super-substrate in vitro and in cells. We solved the structure of SETD2 with bound super-substrate peptide containing a target lysine to methionine mutation, which revealed better interactions involving three of the substituted residues. Our data illustrate that substrate sequence design can strongly increase the activity of protein lysine methyltransferases.

36 MATERIALS SCIENCE↗

Metagenome-Guided Proteomic Quantification of Reductive Dehalogenases in the Dehalococcoides mccartyi -Containing Consortium SDC-9

At groundwater sites contaminated with chlorinated ethenes, fermentable substrates are often added to promote reductive dehalogenation by indigenous or augmented microorganisms. Contemporary bioremediation performance monitoring relies on nucleic acid biomarkers of key organohalide-respiring bacteria, such as Dehalococcoides mccartyi (Dhc). In this work, metagenome sequencing of the commercial, Dhc-containing consortium, SDC-9, identified 12 reductive dehalogenase (RDase) genes, including pceA (two copies), vcrA, and tceA, and allowed for specific detection and quantification of RDase peptides using liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS). Shotgun (i.e., untargeted) proteomics applied to the SDC-9 consortium grown with tetrachloroethene (PCE) and lactate identified 143 RDase peptides, and 36 distinct peptides that covered greater than 99% of the protein-coding sequences of the PceA, TceA, and VcrA RDases. Quantification of RDase peptides using multiple reaction monitoring (MRM) assays with 13 C-/ 15 N-labeled peptides determined 1.8 × 10 3 TceA and 1.2 × 10 2 VcrA RDase molecules per Dhc cell. Ultimately, the MRM mass spectrometry approach allowed for sensitive detection and accurate quantification of relevant Dhc RDases and has potential utility in bioremediation monitoring regimes.

59 BASIC BIOLOGICAL SCIENCES↗

IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data

Proteoforms, the different forms of a protein with sequence variations including post-translational modifications (PTMs), execute vital functions in biological systems such as cell signaling and epigenetic regulation. Precisely defining the stoichiometry of PTMs has been challenging because, in the widely used bottom-up proteomics methods, the detection occurs at the peptide level and thus the link between peptides and their specific modification site is lost, resulting in proteoform ambiguity. Advances in top-down mass spectrometry (MS) technology have permitted the direct characterization of intact proteoforms and their exact number of modification sites, allowing for the relative quantification of positional isomers (PI). Proteins with positional isomers refers to proteoforms with identical total mass and set of modifications but varying PTM site combinations. The relative abundance of PI can be estimated by matching proteoform-specific fragment ions to top-down tandem MS (MS2) data to localize and quantify modifications. However, current approaches heavily rely on manual annotation. Here, we present IsoForma, an open-source R package for relative quantification of PI within a single tool. We benchmarked IsoForma’s performance against two existing workflows and highlight the similarity of the results and improvements in speed. Overall, IsoForma provides a streamlined process, reduces the time of conducting isoform-based analyses, and offers an essential framework for developing customized proteoform analysis workflows. Finally, the software is open source and available at https://github.com/EMSL-Computing/isoforma-lib.

59 BASIC BIOLOGICAL SCIENCES↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗