Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

MarK, a Novosphingobium aromaticivorans kinase required for catabolism of multiple aromatic monomers

The aromatic compounds used in a variety of industrial products are currently obtained from nonrenewable petroleum sources. Alternatively, the plant polymer lignin is an abundant renewable source of aromatics, and its depolymerization generates a variety of products that can include acetovanillone, a vanillin derivative containing an acetyl side chain. The Alphaproteobacterium Novosphingobium aromaticivorans DSM12444 can metabolize several chemically modified aromatics in deconstructed lignin, but not acetovanillone. In this work, adaptive laboratory evolution identified a single amino acid change in the previously uncharacterized gene product Saro_1862 that is necessary and sufficient for N. aromaticivorans growth with acetovanillone as a sole growth substrate, as well as other aromatic monomers not metabolized by wild-type cells. We show that a glutamate (E) to lysine (K) substitution at amino acid residue 16 of Saro_1862 results in a ~1600-fold increase in the rate of ATP-dependent acetovanillone phosphorylation. We also find that recombinant Saro_1862 E16K phosphorylates several other aromatic compounds in vitro , defining the first reported catalytic activity for the widespread UPF0261 protein domain contained in Saro_1862. Thus, we propose naming Saro_1862 MarK, for multiple aromatic kinase. A 1.57 Å crystal structure of MarK E16K predicts that the E16K substitution lies in a potential ATP binding site, suggesting how this amino acid change increased catalytic activity. A search for homologs of MarK and other proteins required for acetovanillone degradation predicts that this pathway for aromatic metabolism exists throughout the bacterial phylogeny.

Novosphingobium↗

Mercury methylation by metabolically versatile and cosmopolitan marine bacteria

Microbes transform aqueous mercury (Hg) into methylmercury (MeHg), a potent neurotoxin that accumulates in terrestrial and marine food webs, with potential impacts on human health. This process requires the gene pair hgcAB, which encodes for proteins that actuate Hg methylation, and has been well described for anoxic environments. However, recent studies report potential MeHg formation in suboxic seawater, although the microorganisms involved remain poorly understood. In this study, we conducted large-scale multi-omic analyses to search for putative microbial Hg methylators along defined redox gradients in Saanich Inlet, British Columbia, a model natural ecosystem with previously measured Hg and MeHg concentration profiles. Analysis of gene expression profiles along the redoxcline identified several putative Hg methylating microbial groups, including Calditrichaeota, SAR324 and Marinimicrobia, with the last the most active based on hgc transcription levels. Marinimicrobia hgc genes were identified from multiple publicly available marine metagenomes, consistent with a potential key role in marine Hg methylation. Computational homology modelling predicts that Marinimicrobia HgcAB proteins contain the highly conserved amino acid sites and folding structures required for functional Hg methylation. Furthermore, a number of terminal oxidases from aerobic respiratory chains were associated with several putative novel Hg methylators. Our findings thus reveal potential novel marine Hg-methylating microorganisms with a greater oxygen tolerance and broader habitat range than previously recognized.

59 BASIC BIOLOGICAL SCIENCES↗

Predicted structural proteome of Sphagnum divinum and proteome-scale annotation

Sphagnum-dominated peatlands store a substantial amount of terrestrial carbon. The genus is undersampled and under-studied. No experimental crystal structure from any Sphagnum species exists in the Protein Data Bank and fewer than 200 Sphagnum-related genes have structural models available in the AlphaFold Protein Structure Database. Tools and resources are needed to help bridge these gaps, and to enable the analysis of other structural proteomes now made possible by accurate structure prediction. We present the predicted structural proteome (25,134 primary transcripts) of Sphagnum divinum computed using AlphaFold, structural alignment results of all high-confidence models against an annotated nonredundant crystallographic database of over 90,000 structures, a structure-based classification of putative Enzyme Commission (EC) numbers across this proteome, and the computational method to perform this proteome-scale structure-based annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Structural analyses of macromolecules by solution scattering (CRADA Final Report)

New innovative Small Angle X-ray Scattering (SAXS) methods to visualized macromolecules in solution at low resolution has led the SIBYLS group at Lawrence Berkeley Lab in being a world leader in biological SAXS. This research is particularly important now as LBNL establishes high-priority research topics including molecular to mesoscale analysis. This new expanding area of research relies on and ties in with our efforts towards developing a solution structure modeling tool for RNA structure prediction. Due to the pandemic our efforts focused on the SARS-CoV-2 proteins and RNA interactions. Specifically, we characterized SARS-CoC-2 proteins interaction with RNA (Wilanowski et al. 2021 -Hammel group) and Nucleoprotein interaction with RNA (Schneidman group).

59 BASIC BIOLOGICAL SCIENCES↗

A chloroplast protein atlas reveals punctate structures and spatial organization of biosynthetic pathways

Chloroplasts are eukaryotic photosynthetic organelles that drive the global carbon cycle. Despite their importance, our understanding of their protein composition, function, and spatial organization remains limited. Here, we determined the localizations of 1,034 candidate chloroplast proteins using fluorescent protein tagging in the model alga Chlamydomonas reinhardtii. The localizations provide insights into the functions of poorly characterized proteins; identify novel components of nucleoids, plastoglobules, and the pyrenoid; and reveal widespread protein targeting to multiple compartments. We discovered and further characterized cellular organizational features, including eleven chloroplast punctate structures, cytosolic crescent structures, and unexpected spatial distributions of enzymes within the chloroplast. We also used machine learning to predict the localizations of other nuclear-encoded Chlamydomonas proteins. The strains and localization atlas developed here will serve as a resource to accelerate studies of chloroplast architecture and functions.

59 BASIC BIOLOGICAL SCIENCES↗

Genome sequence, phylogenetic analysis, and structure-based annotation reveal metabolic potential of Chlorella sp. SLA-04

Algae are a broad class of photosynthetic eukaryotes that are phylogenetically and physiologically diverse. Most of the phylogenetic diversity has been inferred from 18S rDNA sequencing since there are only a few complete genomes available in public databases. Here we use ultra-long-read Nanopore sequencing to determine a gapless, telomere-to-telomere complete genome sequence of Chlorella sp. SLA-04, previously described as Chlorella sorokiniana SLA-04. Chlorella sp. SLA-04 is a green alga that grows to high cell density in a wide variety of environments - high and neutral pH, high and low alkalinity, and high and low salinity. SLA-04's ability to grow in high pH and high alkalinity media without external CO 2 supply is favorable for large-scale algal biomass production. Phylogenetic analysis performed using ribosomal DNA and conserved protein sequences consistently reveal that Chlorella sp. SLA-04 forms a distinct lineage from other strains of Chlorella sorokiniana. We complement traditional genome annotation methods with high throughput structural predictions and demonstrate that this approach expands functional prediction of the SLA-04 proteome. Genomic analysis of the SLA-04 genome identifies the genes capable of utilizing TCA cycle intermediates to replenish cytosolic acetyl-CoA pools for lipid production. We also identify a complete metabolic pathway for sphingolipid anabolism that may allow SLA-04 to readily adapt to changing environmental conditions and facilitate robust cultivation in mass production systems. Altogether, this work clarifies the phylogeny of Chlorella sp. SLA-04 within Trebouxiophyceae and demonstrates how structural predictions can be used to improve annotation beyond sequencebased methods.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-guided design of a broadly cross-reactive multivalent group a streptococcal vaccine

The M protein of group A streptococci (Strep A) is a major virulence determinant and protective antigen. The N-terminal region of the M protein is variable in sequence, defines the M/emm type, and contains epitopes that elicit opsonic antibodies that protect animals from challenge infections. Although there are >200 M types of Strep A, there is now evidence that structurally related M proteins can be grouped into clusters and that immunity may be cluster-specific in addition to M type-specific. This observation has led to recent studies of structure-based design of multivalent M peptide vaccines to select peptides predicted to cross-react with heterologous M types to improve vaccine coverage. In the current study, we have applied a refined series of peptide structural algorithms to predict immunological cross-reactivity among 117N-terminal M peptides representing the most prevalent M types of Strep A. Based on the results of the structural analyses, in combination with global M type prevalence data, we constructed a 32-valent vaccine containing 19 cross-reactive vaccine candidates predicted to cross-react with 37 heterologous M peptides to which were added 13 type-specific M peptides. Further, the 4-protein recombinant vaccine was immunogenic in rabbits and elicited significant levels of antibodies against 31/32 (97%) vaccine peptides and 28/37 (76%) peptides predicted to cross-react. The vaccine antisera also promoted opsonophagocytic killing of vaccine and cross-reactive M types of Strep A. Based on a recent analysis of M type prevalence of Strep A, the potential global coverage of the 32-valent vaccine is ~90%, ranging from 68% in Africa to 95% in North America. Our results indicate the utility of structure-based design that may be applied to future studies of broadly protective M peptide vaccines.

60 APPLIED LIFE SCIENCES↗

Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy

Rational catalyst design remains a significant challenge, with electronic structure, steric, and electrostatic effects known to contribute to activity. Recently, dynamics has been recognized as another factor that impacts catalysis, though identifying and predicting these effects has remained out of reach. Nickel-substituted rubredoxin (NiRd), a protein-based mimic of a hydrogenase enzyme, serves as a model catalytic system in which dynamics can be systematically investigated with respect to activity. While over 30 secondary-sphere mutants of NiRd have been shown to be catalytically active, no significant correlation was observed between the rates and catalytic overpotential or electronic structure, prompting questions about the protein-derived factors that modulate activity. Here, in this work, NMR spectroscopy was used to investigate the roles of substrate accessibility, protein dynamics, and protein stability in controlling catalysis. Significant paramagnetic effects from the nickel center (S = 1) isolate the methylene proton resonances of the metal-coordinating cysteine residues. The sensitivity of resonance positions and linewidths to local environment offers an opportunity to study dynamical molecular changes around the metal center with high resolution. Machine learning algorithms were employed to identify correlations between the catalytic activity and the paramagnetic NMR spectra. These analyses revealed spectroscopic features of specific cysteine protons that report on catalytic overpotential and increased turnover rates, which are further supported by the results obtained using high-field NMR techniques. Collectively, these studies indicate the potential for multifrequency NMR techniques to resolve key contributors to catalytic activity and highlight the importance of local and outer-sphere dynamics.

Protein Engineering↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Fast myosin binding protein C knockout in skeletal muscle alters length-dependent activation and myofilament structure

In striated muscle, the sarcomeric protein myosin-binding protein-C (MyBP-C) is bound to the myosin thick filament and is predicted to stabilize myosin heads in a docked position against the thick filament, which limits crossbridge formation. Here, we use the homozygous Mybpc2 knockout (C2 -/- ) mouse line to remove the fast-isoform MyBP-C from fast skeletal muscle and then conduct mechanical functional studies in parallel with small-angle X-ray diffraction to evaluate the myofilament structure. We report that C2 -/- fibers present deficits in force production and calcium sensitivity. Structurally, passive C2 -/- fibers present altered sarcomere length-independent and -dependent regulation of myosin head conformations, with a shift of myosin heads towards actin. At shorter sarcomere lengths, the thin filament is axially extended in C2 -/- , which we hypothesize is due to increased numbers of low-level crossbridges. These findings provide testable mechanisms to explain the etiology of debilitating diseases associated with MyBP-C.

59 BASIC BIOLOGICAL SCIENCES↗

Persistent Protein Motions in a Rugged Energy Landscape Revealed by Normal Mode Ensemble Analysis

Proteins are allosteric machines that couple motions at distinct, often distant, sites to control biological function. Low-frequency structural vibrations are a mechanism of this long-distance connection and are often used computationally to predict correlations, but experimentally identifying the vibrations associated with specific motions has proved challenging. Spectroscopy is an ideal tool to explore these excitations, but measurements have been largely unable to identify important frequency bands. The result is at odds with some previous calculations and raises the question what methods could successfully characterize protein structural vibrations. Here we show the lack of spectral structure arises in part from the variations in protein structure as the protein samples the energy landscape. However, by averaging over the energy landscape as sampled using an aggregate 18.5 μs of all-atom molecular dynamics simulation of hen egg white lysozyme and normal-mode analyses, we find vibrations with large overlap with functional displacements are surprisingly concentrated in narrow frequency bands. These bands are not apparent in either the ensemble averaged vibrational density of states or isotropic absorption. However, in the case of the ensemble averaged anisotropic absorption, there is persistent spectral structure and overlap between this structure and the functional displacement frequency bands. We systematically lay out heuristics for calculating the spectra robustly, including the need for statistical sampling of the protein and inclusion of adequate water in the spectral calculation. The results show the congested spectrum of these complex molecules obscures important frequency bands associated with function and reveal a method to overcome this congestion by combining structurally sensitive spectroscopy with robust normal mode ensemble analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Strategies to identify and edit improvements in synthetic genome segments episomally

Genome engineering projects often utilize bacterial artificial chromosomes (BACs) to carry multi-kilobase DNA segments at low copy number. However, all stages of whole-genome engineering have the potential to impose mutations on the synthetic genome that can reduce or eliminate the fitness of the final strain. Here, we describe improvements to a multiplex automated genome engineering (MAGE) protocol to improve recombineering frequency and multiplexability. This protocol was applied to recoding an Escherichia coli strain to replace seven codons with synonymous alternatives genome wide. Ten 44 402–47 179 bp de novo synthesized DNA segments contained in a BAC from the recoded strain were unable to complement deletion of the corresponding 33–61 wild-type genes using a single antibiotic resistance marker. Next-generation sequencing (NGS) was used to identify 1–7 non-recoding mutations in essential genes per segment, and MAGE in turn proved a useful strategy to repair these mutations on the recoded segment contained in the BAC when both the recoded and wild-type copies of the mutated genes had to exist by necessity during the repair process. Finally, two web-based tools were used to predict the impact of a subset of non-recoding missense mutations on strain fitness using protein structure and function calls.

59 BASIC BIOLOGICAL SCIENCES↗

Design of Broadly Cross-Reactive M Protein–Based Group A Streptococcal Vaccines

Group A streptococcal infections are a significant cause of global morbidity and mortality. A leading vaccine candidate is the surface M protein, a major virulence determinant and protective Ag. An obstacle to the development of M protein–based vaccines is the >200 different M types defined by the N-terminal sequences that contain protective epitopes. Despite sequence variability, M proteins share coiled-coil structural motifs that bind host proteins required for virulence. In this study, we exploit this potential Achilles heel of conserved structure to predict cross-reactive M peptides that could serve as broadly protective vaccine Ags. Combining sequences with structural predictions, six heterologous M peptides in a sequence-related cluster were predicted to elicit cross-reactive Abs with the remaining five nonvaccine M types in the cluster. The six-valent vaccine elicited Abs in rabbits that reacted with all 11 M peptides in the cluster and functional opsonic Abs against vaccine and nonvaccine M types in the cluster. We next immunized mice with four sequence-unrelated M peptides predicted to contain different coiled-coil propensities and tested the antisera for cross-reactivity against 41 heterologous M peptides. Based on these results, we developed an improved algorithm to select cross-reactive peptide pairs using additional parameters of coiled-coil length and propensity. The revised algorithm accurately predicted cross-reactive Ab binding, improving the Matthews correlation coefficient from 0.42 to 0.74. These results form the basis for selecting the minimum number of N-terminal M peptides to include in potentially broadly efficacious multivalent vaccines that could impact the overall global burden of group A streptococcal diseases.

60 APPLIED LIFE SCIENCES↗

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of a secretory heme‐binding protein from Nocardia seriolae involved in cell apoptosis

Abstract According to the whole‐genome bioinformatics analysis, a heme‐binding protein from Nocardia seriolae (HBP) was found. HBP was predicted to be a bacterial secretory protein, located at mitochondrial membrane in eukaryotic cells and have a similar protein structure with the heme‐binding protein of Mycobacterium tuberculosis , Rv0203. In this study, HBP was found to be a secretory protein and co‐localized with mitochondria in FHM cells. Quantitative analysis of mitochondrial membrane potential value, caspase‐3 activity, and transcription level of apoptosis‐related genes suggested that overexpression of HBP protein can induce cell apoptosis. In conclusion, HBP was a secretory protein which may target to mitochondria and involve in cell apoptosis in host cells. This research will promote the function study of HBP and deepen the comprehension of the virulence factors and pathogenic mechanisms of N. seriolae .

Wen, Yiming↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗