Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Sequences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Engineering Pseudomonas putida for production of 3-hydroxyacids using hybrid type I polyketide synthases

Engineered type I polyketide synthases (T1PKSs) are a potentially transformative platform for the biosynthesis of small molecules. Due to their modular nature, T1PKSs can be rationally designed to produce a wide range of bulk or specialty chemicals. While heterologous PKS expression is best studied in microbes of the genus Streptomyces, recent studies have focused on the exploration of non-native PKS hosts. The biotechnological production of chemicals in fast growing and industrial relevant hosts has numerous economic and logistic advantages. With its native ability to utilize alternative feedstocks, Pseudomonas putida has emerged as a promising workhorse for the sustainable production of small molecules. Here, we outline the assessment of P. putida as a host for the expression of engineered T1PKSs and production of 3-hydroxyacids. After establishing the functional expression of an engineered T1PKS, we successfully expanded and increased the pool of available acyl-CoAs needed for the synthesis of polyketides using transposon sequencing and protein degradation tagging. This work demonstrates the potential of T1PKSs in P. putida as a production platform for the sustainable biosynthesis of unnatural polyketides.

Schmidt, Matthias↗

Directing Nanoparticle Organization in Response to Diverse Chemical Inputs

Signaling cascades are crucial for transducing stimuli in biological systems, enabling multiple stimuli to regulate a downstream target with precisely controlled timing and amplifying signals through a series of intermediary reactions. Developing a robust signaling system with such capabilities would be pivotal for programming complex behaviors in synthetic DNA-based molecular devices. However, although “software” such as nucleic acid circuits could potentially be harnessed to relay signals to DNA-based nanostructure hardware, such explorations have been limited. Here, in this study, we develop a platform for transducing a variety of stimuli via messenger-mediated reactions to regulate the release and reloading of gold nanoparticles (AuNPs) in a 3D DNA framework. In the first step, an in vitro transcription circuit is engineered to sense and amplify chemical stimuli, including arbitrary DNA sequences and proteins, producing RNA. In the second step, the RNA releases the DNA-coated AuNPs from the DNA framework via a strand displacement reaction. AuNP reloading is controlled by a separate step driven by degradation of the RNA. Our platform holds promise for applications requiring dynamic multiagent control over DNA-based devices, offering a versatile tool for advanced molecular device engineering.

36 MATERIALS SCIENCE↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES↗

Frameshifting Stimulatory Sequence Induces Large Structural Change of Ribosomal Proteins When Bound to E. coli Ribosomes

Biological macromolecular machines occupy a continuum of structural conformations to perform cellular tasks. Mapping this conformational space provides an insight into its functionality. While the cryo-electron microscopy resolution revolution has expanded our ability to characterize the conformational continuums, there are obstacles in structurally characterizing regions of high flexibility. These technical barriers have impeded characterization of flexible ribosomal proteins when the ribosome is interacting with mRNA stem-loop structures such as a frameshifting stimulatory sequence (FSS). Small-angle neutron/X-ray scattering and electron microscopy were used to study ribosomal samples and compared structural differences between a ribosome that is bound to an FSS stem-loop compared to a ribosome bound to linear mRNA. This comparison shows that a large protein stalk elongates by 22% when the 70S interacts with an mRNA stem-loop. Finally, our results suggest that ribosomal proteins have extensive flexibility and may influence important ribosomal mechanisms, such as those that involve FSS.

36 MATERIALS SCIENCE↗

Water, Solute, and Ion Transport in De Novo-Designed Membrane Protein Channels

Biological organisms engineer peptide sequences to fold into membrane pore proteins capable of performing a wide variety of transport functions. Synthetic de novo-designed membrane pores can mimic this approach to achieve a potentially even larger set of functions. Here, in this work, we explore water, solute, and ion transport in three de novo designed β-barrel membrane channels in the 5–10 Å pore size range. We show that these proteins form passive membrane pores with high water transport efficiencies and size rejection characteristics consistent with the pore size encoded in the protein structure. Ion conductance and ion selectivity measurements also show trends consistent with the pore size, with the two larger pores showing weak cation selectivity. MD simulations of water and ion transport and solute size exclusion are consistent with the experimental trends and provide further insights into structure–function correlations in these membrane pores.

59 BASIC BIOLOGICAL SCIENCES↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

Designing Peptide Fossils That Model the Evolution of the Bacterial Ferredoxin Fold

Electron transfer coupled to redox chemistry is at the heart of metabolism. The proteins responsible for moving electrons (protein electron carriers) must have emerged at the origin of life. The small iron–sulfur-binding bacterial ferredoxins were likely among these first proteins. Embedded within the ferredoxin sequence and structure is a symmetry that points to an ancient gene duplication event. Little is understood about the nature of ferredoxins prior to this duplication event or what environmental factors may have driven the selection for more complex forms. The deep-time molecular history of ferredoxins goes back billions of years and cannot be reconstructed by phylogenetic analyses based on amino acid sequences. Here, we use structure-guided protein design to model a fossil half-ferredoxin stage in the evolution of this fold, the semidoxins, and their symmetric full-length counterparts, the symdoxins. Semidoxin designs homodimerize, exhibiting structural, thermodynamic, and electrochemical behaviors in most cases identical to cognate symdoxins. However, the semi- and symdoxin fossil stages behave differently when incorporated into an in vivo electron transfer complementation assay. Both can support bacterial growth dependent on protein expression. Growth rates of bacteria expressing the semidoxins are much more sensitive to oxygen than those of bacteria expressing symdoxins. Motivated by the in vivo functionality of designed semidoxins, we identified putative naturally occurring semidoxins in extant anaerobic microorganisms. This is consistent with the observed in vivo oxygen sensitivity of the semidoxin designs. One natural semidoxin is shown to be folded and redox active. However, it exists as a mixture of monomers and dimers, suggesting a potential connection between semidoxins and even simpler single iron–sulfur cluster-binding peptides.

59 BASIC BIOLOGICAL SCIENCES↗

Blocking C-terminal processing of KRAS4b via a direct covalent attack on the CaaX-box cysteine

RAS is the most frequently mutated oncogene in cancer. RAS proteins show high sequence similarities in their G-domains but are significantly different in their C-terminal hypervariable regions (HVR). These regions interact with the cell membrane via lipid anchors that result from posttranslational modifications (PTM) of cysteine residues. KRAS4b is unique as it has only one cysteine that undergoes PTM, C185. Small molecule covalent modification of C185 would block any form of prenylation and subsequently inhibit attachment of KRAS4b to the cell membrane, blocking its biological activity. We translated this concept to the discovery and development of disulfide tethering screen hits into irreversible covalent modifiers of C185. These compounds inhibited proliferation of KRAS4b-driven mouse embryonic fibroblasts, but not cells driven by N-myristoylated KRAS4b that harbor a C185S mutation and are not dependent on C185 prenylation. Top–down proteomics was used to confirm target engagement in cells. These compounds bind in a pocket formed when the HVR folds back between helix 3 and 4 in the G-domain (HVR-α3-α4). This interaction can happen in the absence of small molecules as predicted by molecular dynamics simulations and is stabilized in the presence of C185 binders as confirmed by small-angle X-ray scattering and solution NMR. NOESY-HSQC, an NMR approach that measures internuclear distances of 6 Å or less, and structure analysis identified the critical residues and interactions that define the HVR-α3-α4 pocket. Further development of compounds that bind to this pocket could be the basis of a new approach to targeting KRAS cancers.

C185↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗