Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Sequences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

Different chemical scaffolds bind to L-phe site in Mycobacterium tuberculosis Phe-tRNA synthetase

Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mt), is one of the deadliest infectious diseases. The rise of multidrug-resistant strains represents a major public health threat, requiring new therapeutic options. Bacterial aminoacyl-tRNA synthetases (aaRS) have been shown to be highly promising drug targets, including for TB treatment. These enzymes play an essential role in translating the DNA gene code into protein sequence by attaching specific amino acid to their cognate tRNAs. They have multiple binding sites that can be targeted for inhibitor discovery: amino acid binding pocket, ATP binding pocket, tRNA binding site and an editing domain. Recently we reported several high-resolution structures of M. tuberculosis phenylalanyl-tRNA synthetase (MtPheRS) complexed with tRNA Phe and either L-Phe or a nonhydrolyzable phenylalanine adenylate analog. Here, in this study, using Nucleic Magnetic Resonance (NMR) and Surface Plasmon Resonance (SPR) we identified fragments that bind to MtPheRS and we determined crystal structures of their complexes with MtPheRS/tRNA Phe . All the binders interact with the L-Phe amino acid binding site. The analysis of interactions of the new compounds combined with adenylate analog structure provides insights for the rational design of antituberculosis drugs. The 3 ' arm of the tRNA Phe in all the structures was disordered with exception of one complex with D-735 compound. In this structure the 3' CCA end of the acceptor stem is observed in the editing domain of MtPheRS providing insights regarding the post-transfer editing activity of class II aaRS.

Gade, Priyanka [Univ. of Chicago, IL (United State↗

Evolutionary trajectory of transcription factors and selection of targets for metabolic engineering

Transcription factors (TFs) provide potentially powerful tools for plant metabolic engineering as they often control multiple genes in a metabolic pathway. However, selecting the best TF for a particular pathway has been challenging, and the selection often relies significantly on phylogenetic relationships. Here, we offer examples where evolutionary relationships have facilitated the selection of the suitable TFs, alongside situations where such relationships are misleading from the perspective of metabolic engineering. We argue that the evolutionary trajectory of a particular TF might be a better indicator than protein sequence homology alone in helping decide the best targets for plant metabolic engineering efforts. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

ProteinTuneRL

ProteinTuneRL is a framework designed to harness the power of reinforcement learning for advanced protein design. The project enables fine-tuning of generative models to explore and optimize protein sequences with tailored structural and functional properties.

Landajuela Larma, Mikel [Lawrence Livermore Nation↗

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo↗

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]↗

Biosynthesis of bioprivileged, linear molecules via novel carboligase reactions

Over the award period, we made progress on the three aims. We screened twenty-five carboligases for activity coupling twenty-one possible -keto acids (Aim 1). The carboligases were selected across a diverse set of protein sequences. Using Q-Exactive UHPLC-MS, we tested a total of 210 coupled products per enzyme and generated a dataset of 5250 enzyme-substrate activity relationships. We identified multiple enzymes that had activity for synthesizing suberic acid and heptanoic acid (Aim 2). We built a random forest model for predicting the activity of each enzyme toward substrates on which it was not tested using the data from Aim 1. Finally, we evaluated growth defects that occurred due to expression of different carboligases in E. coli (Aim 3). We were able to identify specific metabolites and putative pathways that, when supplemented in the media, recovered the growth defect associated with the presence of specific carboligases. We are in the process of publishing two manuscript describing the methods for high-throughput screening of enzyme promiscuity, using machine learning to predict activity on untested substrates, and enzyme activity data we collected. This project has produced enabling data for biosynthesis of a range of new-to-nature compounds to support biomanufacturing.

60 APPLIED LIFE SCIENCES↗

Rational Design of Lanmodulin Variants for Size-Based Selectivity of Individual Rare Earth Elements

Rare earth elements (REEs) are essential to modern technologies, yet their high physical and chemical similarity makes separation of individual REEs difficult and environmentally taxing. Metalloproteins offer a promising alternative for selective REE binding, as they tend to have high metal ion affinity and specificity. Lanmodulin (LanM), in particular, has arisen as a potential candidate for REE separation as it exhibits picomolar affinity for elements in the REE family. Prior work has shown that the single point mutation D9N can shift LanM’s preference away from lanthanides toward actinides, motivating efforts to tune selectivity of LanM through targeted mutagenesis. Here, we tested the hypothesis that introducing selective aspartic acid to glutamic acid substitutions in the metal coordinating EF hands of LanM would impose steric constraints that would drive LanM affinity away from larger ions, such as La3+, to smaller ions, such as Y3+. To test this hypothesis, a combination of computational and experimental approaches were employed to evaluate the signal mutations LanM D5E and LanM D3E and the double mutants LanM D1ED5E and LanM D3ED9E. Surprisingly, increasing the number of mutations within the metal center did not enhance affinity for smaller REEs, or decrease affinity for larger ions. Only the single point mutation LanM D5E weakened La3+ binding by one order of magnitude relative to LanM wild type (WT), and pairing it with a second mutation to produce LanM D1ED5E drove La3+ affinity to be stronger than that seen for LanM WT. The D3E mutation alone prevented proper expression and folding, but paring it with D9E to produce LanM D3ED9E rescued expression and yielded La3+ affinities comparable to LanM WT. All variants that expressed (LanM D5E, LanM D1ED5E, LanM D3ED9E) displayed Y3+ affinities comparable to LanM WT. Overall, these results highlight the tunability of LanM’s metal-binding environment but also expose current limitations in predicting structural responses to point mutations within a protein sequence. This work establishes a foundation that can be used for refining computational and experimental strategies to engineer metalloproteins with tailored REE selectivity.

Close, Emily [Pacific Northwest National Laborator↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

Signal sequences target enzymes and structural proteins to bacterial microcompartments and are critical for microcompartment formation

ABSTRACT Spatial organization of pathway enzymes has emerged as a promising tool to address several challenges in metabolic engineering, such as flux imbalances and off-target product formation. Bacterial microcompartments (MCPs) are a spatial organization strategy used natively by many bacteria to encapsulate metabolic pathways that produce toxic, volatile intermediates. Several recent studies have focused on engineering MCPs to encapsulate heterologous pathways of interest, but how this engineering affects MCP assembly and function is poorly understood. In this study, we investigated the role of signal sequences, short domains that target proteins to the MCP core, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterized two novel Pdu signal sequences on the structural proteins PduM and PduB, which constitute the first report of metabolosome signal sequences on structural proteins rather than enzymes. We then explored the role of enzymatic and structural Pdu signal sequences on MCP assembly by deleting their encoding sequences from the genome alone and in combination. Deleting enzymatic signal sequences decreased the MCP formation, but this defect could be recovered in some cases by overexpressing genes encoding the knocked-out signal sequence fused to a heterologous protein. By contrast, deleting structural signal sequences caused similar defects to knocking out the genes encoding the full-length PduM and PduB proteins. Our results contribute to a growing understanding of how MCPs form and function in bacteria and provide strategies to mitigate assembly disruption when encapsulating heterologous pathways in MCPs. IMPORTANCE Spatially organizing biosynthetic pathway enzymes is a promising strategy to increase pathway throughput and yield. Bacterial microcompartments (MCPs) are proteinaceous organelles that many bacteria natively use as a spatial organization strategy to encapsulate niche metabolic pathways, providing significant metabolic benefits. Encapsulating heterologous pathways of interest in MCPs could confer these benefits to industrially relevant pathways. Here, we investigate the role of signal sequences, short domains that target proteins for encapsulation in MCPs, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterize two novel signal sequences on structural proteins, constituting the first Pdu signal sequences found on structural proteins rather than enzymes, and perform knockout studies to compare the impacts of enzymatic and structural signal sequences on MCP assembly. Our results demonstrate that enzymatic and structural signal sequences play critical but distinct roles in Pdu MCP assembly and provide design rules for engineering MCPs while minimizing disruption to MCP assembly.

Johnson, Elizabeth R. (ORCID:0000000179236881)↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗