Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

Specialization Restricts the Evolutionary Paths Available to Yeast Sugar Transporters

Functional innovation at the protein level is a key source of evolutionary novelties. The constraints on functional innovations are likely to be highly specific in different proteins, which are shaped by their unique histories and the extent of global epistasis that arises from their structures and biochemistries. These contextual nuances in the sequence–function relationship have implications both for a basic understanding of the evolutionary process and for engineering proteins with desirable properties. Here, we have investigated the molecular basis of novel function in a model member of an ancient, conserved, and biotechnologically relevant protein family. These Major Facilitator Superfamily sugar porters are a functionally diverse group of proteins that are thought to be highly plastic and evolvable. By dissecting a recent evolutionary innovation in an α-glucoside transporter from the yeast Saccharomyces eubayanus, we show that the ability to transport a novel substrate requires high-order interactions between many protein regions and numerous specific residues proximal to the transport channel. To reconcile the functional diversity of this family with the constrained evolution of this model protein, we generated new, state-of-the-art genome annotations for 332 Saccharomycotina yeast species spanning ~400 My of evolution. By integrating phylogenetic and phenotypic analyses across these species, we show that the model yeast α-glucoside transporters likely evolved from a multifunctional ancestor and became subfunctionalized. The accumulation of additive and epistatic substitutions likely entrenched this subfunction, which made the simultaneous acquisition of multiple interacting substitutions the only reasonably accessible path to novelty.

59 BASIC BIOLOGICAL SCIENCES↗

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

The Gene Ontology knowledgebase in 2026

Abstract The Gene Ontology (GO) knowledgebase (https://geneontology.org) is a comprehensive resource describing the functions of genes. The GO knowledgebase is regularly updated and improved. We describe here the major updates that have been made in the past 3 years. The ontology and annotations have been expanded and revised, particularly in several areas of biology: cellular metabolism, multi-organism interactions (e.g. host-pathogen), extracellular matrix proteins, chromatin remodeling (e.g. the “histone code”), and noncoding RNA functions. We have released version 2 of a comprehensive set of integrated, reviewed annotations for human genes, which we call the “functionome.” We have also dramatically increased the number of GO-CAM models, with over 1500 models of metabolic and signaling pathways, primarily in human, mouse, budding and fission yeast, and fruit fly. Finally, we discuss our current recommendations and future prospects of AI in the use and development of GO.

Aleksander, Suzi A (ORCID:0000000167872901)↗

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato↗

Shedding Light on Microbial Dark Matter with A Universal Language of Life

The majority of microbial genomes have yet to be cultured, and most proteins predicted from microbial genomes or sequenced from the environment cannot be functionally annotated. As a result, current computational approaches to describe microbial systems rely on incomplete reference databases that cannot adequately capture the full functional diversity of the microbial tree of life, limiting our ability to model high-level features of biological sequences. The scientific community needs a means to capture the functionally and evolutionarily relevant features underlying biology, independent of our incomplete reference databases. Such a model can form the basis for transfer learning tasks, enabling downstream applications in environmental microbiology, medicine, and bioengineering. Here we present LookingGlass, a deep learning model capturing a “universal language of life”. LookingGlass encodes contextually-aware, functionally and evolutionarily relevant representations of short DNA reads, distinguishing reads of disparate function, homology, and environmental origin. We demonstrate the ability of LookingGlass to be fine-tuned to perform a range of diverse tasks: to identify novel oxidoreductases, to predict enzyme optimal temperature, and to recognize the reading frames of DNA sequence fragments. LookingGlass is the first contextually-aware, general purpose pre-trained “biological language” representation model for short-read DNA sequences. LookingGlass enables functionally relevant representations of otherwise unknown and unannotated sequences, shedding light on the microbial dark matter that dominates life on Earth.

A Hoarfrost↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Characterization of the biofilm landscape of Bacillus subtilis by spatial microproteomics

Bulk proteomics has been demonstrated to differentiate subpopulations within bacterial colonies, yet advanced analyses by mass spectrometry imaging (MSI) hold even greater promise for the future. This technology can enable high-throughput spatial phenotyping that can reshape biological discovery by providing visualization of components of various biomolecular mechanisms. With high mass resolving power and high spatial resolution analyses being routine, we can confidently enable intact protein imaging directly from samples with minimal preparation. Pairing those analyses with bulk experimental libraries can provide high confidence in annotations of post-translational modifications (PTMs) and truncations. Revealing PTM localization within the samples unlocks a direct window into unknown biology at the microscale. However, top-down proteomics (TDP) is not commonplace for microbial species, largely due to challenges in identifying detected peptides and proteins; considering the theoretical proteome of even the well-studied model bacterium Bacillus subtilis was only partially mapped recently. With little still known about the form and function of many of these proteins – let alone proteoforms, where PTMs and truncations of the same protein may possess unique physiological roles – there is a wealth of work to be done. Here we jointly apply TDP and MSI to describe the microscale spatial proteomic landscape within B. subtilis and further demonstrate the feasibility of detecting differentiated subpopulations through proteoforms across the biofilm landscape.

bacterial biofilms↗

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Enzyme property prediction using artificial intelligence

Artificial intelligence (AI)-driven enzyme property prediction enables rapid discovery and engineering of enzymes for a wide range of biotechnological and therapeutic applications. Here, we first introduce the key components in AI model development, including enzyme datasets, protein representation methods, and model architectures. We then highlight a variety of AI tools developed for the prediction of enzyme properties and functional annotations, including enzyme structure, kinetic parameters, substrate specificity, thermostability, solubility, Enzyme Commission number, and Gene Ontology term. Moreover, we describe representative downstream applications enabled by these AI tools. Finally, we discuss some challenges and opportunities as well as future prospects.

Yuan, Le [University of Illinois at Urbana-Champai↗

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES↗

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G↗