Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Sequences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation

This dataset accompanies the publication "ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation". This paper introduces a new AI model for protein sequence generation. This dataset contains data related to experiments discussed in the publication. This includes generated sequences and evaluation metrics supporting all unconditional and bias-controlled experiments in the ProtNHF paper.

60 APPLIED LIFE SCIENCES↗

Emerging protein sequencing technologies: proteomics without mass spectrometry?

Liquid chromatography-tandem mass spectrometry (LC-MS/MS) has been a leading method for proteomics for 30 years. Advantages provided by LC-MS/MS are offset by significant disadvantages, including cost. Recently, several non-mass spectrometric methods have emerged, but little information is available about their capacity to analyze the complex mixtures routine for mass spectrometry. Areas Covered: We review recent non-mass-spectrometric methods for sequencing proteins and peptides, including those using nanopores, sequencing by degradation, reverse translation, and short-epitope mapping, with comments on bioinformatics challenges, fundamental limitations, and areas where new technologies will be more or less competitive with LC-MS/MS. In addition to conventional literature searches, instrument vendor websites, patents, webinars, and preprints were also consulted to give a more up-to-date picture. Expert Opinion: Many new technologies are promising. However, demonstrations that they outperform mass spectrometry in terms of peptides and proteins identified have not yet been published, and astute observers note important disadvantages, especially relating to the dynamic range of single-molecule measurements of complex mixtures. Still, even if the performance of emerging methods proves inferior to LC-MS/MS, their low cost could create a different kind of revolution: a dramatic increase in the number of biology laboratories engaging in new forms of proteomics research.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Dissecting neurofilament tail sequence-phosphorylation-structure relationships with multicomponent reconstituted protein brushes

Neurofilaments (NFs) are multisubunit, bottlebrush-shaped intermediate filaments abundant in the axonal cytoskeleton. Each NF subunit contains a long intrinsically disordered tail domain, which protrudes from the NF core to form a “brush” surrounding each NF. Precisely how the tails’ variable charge patterns and repetitive phosphorylation sites mediate their conformation within the brush remains an open question in axonal biology. We address this problem by grafting recombinant NF tail protein constructs NF-Light, -Medium, and -Heavy (NFL, NFM, and NFH) to surfaces, yielding protein brushes of defined stoichiometry that can be phosphorylated in vitro. Atomic force microscopy measurements reveal that brush height depends on composition monotonically but not always linearly for binary NFL:NFM or NFL:NFH systems, and that NFM-based brushes are highly extended, while brushes incorporating the much larger NFH are surprisingly compact even after multisite phosphorylation. Complementary self-consistent field theory (SCFT) predicts multilayer brush morphologies for NFM and phosphorylated NFH brushes. Further experiments and SCFT analysis with designed mutants reveal that N-terminal negative charges in the NFH tail repel phosphorylated residues to generate the multilayer morphology, while the C-terminal charge-neutral region contributes to multilayer brush morphology but not total brush height. Charge-shuffled NFM variants show that charge segregation promotes brush collapse near physiological ionic strengths. Collectively, this study supports a role for NFM in establishing a dynamic range for NF brush conformation, lending insight into previous in vitro and in vivo findings. More broadly, this work establishes a platform for dissecting contributions of disordered protein sequence to conformation at interfaces.

Science & Technology - Other Topics↗

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto↗

Sequence-defined structural transitions by calcium-responsive proteins

Biopolymer sequences dictate their functions, and protein-based polymers are a promising platform to establish sequence–function relationships for novel biopolymers. To efficiently explore vast sequence spaces of natural proteins, sequence repetition is a common strategy to tune and amplify specific functions. This strategy is applied to repeats-in-toxin (RTX) proteins with calcium-responsive folding behavior, which stems from tandem repeats of the nonapeptide GGXGXDXUX in which X can be any amino acid and U is a hydrophobic amino acid. To determine the functional range of this nonapeptide, we modified a naturally occurring RTX protein that forms β-roll structures in the presence of calcium. Sequence modifications focused on calcium-binding turns within the repetitive region, including either global substitution of nonconserved residues or complete replacement with tandem repeats of a consensus nonapeptide GGAGXDTLY. Some sequence modifications disrupted the typical transition from intrinsically disordered random coils to folded β rolls, despite conservation of the underlying nonapeptide sequence. Proteins enriched with smaller, hydrophobic amino acids adopted secondary structures in the absence of calcium and underwent structural rearrangements in calcium-rich environments. In contrast, proteins with bulkier, hydrophilic amino acids maintained intrinsic disorder in the absence of calcium. In conclusion, these results indicate a significant role of nonconserved amino acids in calcium-responsive folding, thereby revealing a strategy to leverage sequences in the design of tunable, calcium-responsive biopolymers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,↗

Uncovering Sequence and Structural Characteristics of Fungal Expansin‐Related Proteins With Potential to Drive Substrate Targeting

Expansins loosen plant cell wall networks through disrupting non-covalent bonds between cellulose microfibrils and matrix polysaccharides. Whereas expansins were first discovered in plants, expansin-related proteins have since been identified in bacteria and fungi. The biological function of microbial expansins remains unclear; however, several studies have shown distinct binding preferences toward different structural polysaccharides. Earlier studies of bacterial expansin-related proteins uncovered sequence and structural features that correlate to substrate binding. Herein, 20 fungal expansin-related sequences were recombinantly produced in Komagataella phaffii, and the purified proteins were compared in terms of substrate binding to cellulosic and chitinous substrates. The impact of pH on the zeta potential of prioritized substrates was also measured, and Principal Component Analysis was performed to uncover correlations between protein characteristics (e.g., pI, hydrophobicity, surface charge distribution) and measured substrate binding preferences. Whereas acidic proteins with a predicted pI less than 5.0 preferentially bound to chitin, basic proteins with pI greater than 8.0 preferentially bound to xylan and xylan-containing fiber. Similar to many cellulases, binding to cellulose was correlated to relatively high aromatic amino acid content in the protein sequence and presence of a carbohydrate binding module (CBM), which in the case of expansins is a C-terminal CBM63. Whereas overall sequence characteristics could be correlated to substrate binding preference, the identity of amino acids occupying conserved positions that impact protein activity was better correlated with loosenin versus expansin classifications.

chitin↗

Sensitive and error-tolerant annotation of protein-coding DNA with BATH

We present BATH, a tool for highly sensitive annotation of protein-coding DNA based on direct alignment of that DNA to a database of protein sequences or profile hidden Markov models (pHMMs). BATH is built on top of the HMMER3 code base, and simplifies the annotation workflow for pHMM-based translated sequence annotation by providing a straightforward input interface and easy-to-interpret output. BATH also introduces novel frameshift-aware algorithms to detect frameshift-inducing nucleotide insertions and deletions (indels). BATH matches the accuracy of HMMER3 for annotation of sequences containing no errors, and produces superior accuracy to all tested tools for annotation of sequences containing nucleotide indels. These results suggest that BATH should be used when high annotation sensitivity is required, particularly when frameshift errors are expected to interrupt protein-coding regions, as is true with long-read sequencing data and in the context of pseudogenes.

59 BASIC BIOLOGICAL SCIENCES↗

Structural genomics of bacterial drug targets: Application of a high-throughput pipeline to solve 58 protein structures from pathogenic and related bacteria

Antibiotic resistance remains a leading cause of severe infections worldwide. Small changes in protein sequence can impact antibiotic efficacy. Here, we report deposition of 58 X-ray crystal structures of bacterial proteins that are known targets for antibiotics, which expands knowledge of structural variation to support future antibiotic discovery or modifications.

PDB↗

Characterizing Integrase-Attachment Site Pairs: A Machine Learning Approach

Genomic islands (GIs) are mobile genetic elements that integrate into host genomes via self-encoded integrases at specific DNA sequences known as attachment (att) sites. The ability to predict the target att site from an integrase's protein sequence is a central challenge in genomics due to the sequence diversity of integrases and the subtlety of their DNA recognition motifs.

59 BASIC BIOLOGICAL SCIENCES↗

A generalized platform for artificial intelligence-powered autonomous enzyme engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-fold improvement in substrate preference and 16-fold improvement in ethyltransferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

Strangers in a foreign land: ‘Yeastizing’ plant enzymes

Abstract Expressing plant metabolic pathways in microbial platforms is an efficient, cost‐effective solution for producing many desired plant compounds. As eukaryotic organisms, yeasts are often the preferred platform. However, expression of plant enzymes in a yeast frequently leads to failure because the enzymes are poorly adapted to the foreign yeast cellular environment. Here, we first summarize the current engineering approaches for optimizing performance of plant enzymes in yeast. A critical limitation of these approaches is that they are labour‐intensive and must be customized for each individual enzyme, which significantly hinders the establishment of plant pathways in cellular factories. In response to this challenge, we propose the development of a cost‐effective computational pipeline to redesign plant enzymes for better adaptation to the yeast cellular milieu. This proposition is underpinned by compelling evidence that plant and yeast enzymes exhibit distinct sequence features that are generalizable across enzyme families. Consequently, we introduce a data‐driven machine learning framework designed to extract ‘yeastizing’ rules from natural protein sequence variations, which can be broadly applied to all enzymes. Additionally, we discuss the potential to integrate the machine learning model into a full design‐build‐test cycle.

59 BASIC BIOLOGICAL SCIENCES↗

Engineered Membrane Vesicle Production via oprF or oprI Deletion Has Distinct Phenotypic Effects in Pseudomonas putida -putative knockouts table

Table S1, putative gene knockout targets in P. putida KT2440 to enhance vesiculation; Table S2, protein sequence identity of OmpA from E. coli K12 to P. putida KT2440 genes; Table S3, strains utilized in this study and corresponding construction details; Table S4, oligonucleotides utilized in this study; Table S5, plasmids utilized in this study; Table S6, sequences for mNeonGreen, tags, and codon-optimized genes; Figure S1, particle count per gCDW for KT2440 and knockout strains corresponding to data presented in Figure 1B; Figure S2, OD600 measurements of extracted MVs from KT2440 and knockout strains; Figure S3, particle count per gCDW for WT, ΔPP_4669, and ΔPP_1502; Figure S4, particle count per gCDW for KT2440 and knockout strains corresponding to data presented in Figure 3C; Figure S5, sizes of MVs corresponding to particle counts in Figure S4; Figure S6, particle count per gCDW for KT2440 grown on 20 mM glucose alone or 20 mM glucose plus 12.5 mM p-coumarate and 12.5 mM ferulate; Figure S7 and Figure S8, principal component analysis of the cellular fractions; Figure S9, heatmap of outer membrane proteins with differential abundance; and Figure S10, mNeonGreen (mNG) fluorescence signal for the cellular fraction and the extracellular fraction

hypervesiculation↗