Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Evaluating the Impact of gRNA SNPs in CasRx Activity for Reducing Viral RNA in HCoV-OC43

Viruses within a given family often share common essential genes that are highly conserved due to their critical role for the virus’s replication and survival. In this work, we developed a proof-of-concept for a pan-coronavirus CRISPR effector system by designing CRISPR targets that are cross-reactive among essential genes of different human coronaviruses (HCoV). We designed CRISPR targets for both the RNA-dependent RNA polymerase (RdRp) gene as well as the nucleocapsid (N) gene in coronaviruses. Using sequencing alignment, we determined the most highly conserved regions of these genes to design guide RNA (gRNA) sequences. In regions that were not completely homologous among HCoV species, we introduced mismatches into the gRNA sequence and tested the efficacy of CasRx, a Cas13d type CRISPR effector, using reverse transcription quantitative polymerase chain reaction (RT-qPCR) in HCoV-OC43. We evaluated the effect that mismatches in gRNA sequences has on the cleavage activity of CasRx and found that this CRISPR effector can tolerate up to three mismatches while still maintaining its nuclease activity in HCoV-OC43 viral RNA. This work highlights the need to evaluate off-target effects of CasRx with gRNAs containing up to three mismatches in order to design safe and effective CRISPR experiments.

59 BASIC BIOLOGICAL SCIENCES↗

A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers

Abstract Motivation Deep learning has revolutionized protein tertiary structure prediction recently. The cutting-edge deep learning methods such as AlphaFold can predict high-accuracy tertiary structures for most individual protein chains. However, the accuracy of predicting quaternary structures of protein complexes consisting of multiple chains is still relatively low due to lack of advanced deep learning methods in the field. Because interchain residue–residue contacts can be used as distance restraints to guide quaternary structure modeling, here we develop a deep dilated convolutional residual network method (DRCon) to predict interchain residue–residue contacts in homodimers from residue–residue co-evolutionary signals derived from multiple sequence alignments of monomers, intrachain residue–residue contacts of monomers extracted from true/predicted tertiary structures or predicted by deep learning, and other sequence and structural features. Results Tested on three homodimer test datasets (Homo_std dataset, DeepHomo dataset and CASP-CAPRI dataset), the precision of DRCon for top L/5 interchain contact predictions (L: length of monomer in a homodimer) is 43.46%, 47.10% and 33.50% respectively at 6 Å contact threshold, which is substantially better than DeepHomo and DNCON2_inter and similar to Glinter. Moreover, our experiments demonstrate that using predicted tertiary structure or intrachain contacts of monomers in the unbound state as input, DRCon still performs well, even though its accuracy is lower than using true tertiary structures in the bound state are used as input. Finally, our case study shows that good interchain contact predictions can be used to build high-accuracy quaternary structure models of homodimers. Availability and implementation The source code of DRCon is available at https://github.com/jianlin-cheng/DRCon. The datasets are available at https://zenodo.org/record/5998532#.YgF70vXMKsB. Supplementary information Supplementary data are available at Bioinformatics online.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sequence-structure-function characterization of the emerging tetracycline destructase family of antibiotic resistance enzymes

Tetracycline destructases (TDases) are flavin monooxygenases which can confer resistance to all generations of tetracycline antibiotics. The recent increase in the number and diversity of reported TDase sequences enables a deep investigation of the TDase sequence-structure-function landscape. Here, we evaluate the sequence determinants of TDase function through two complementary approaches: (1) constructing profile hidden Markov models to predict new TDases, and (2) using multiple sequence alignments to identify conserved positions important to protein function. Using the HMM-based approach we screened 50 high-scoring candidate sequences in Escherichia coli, leading to the discovery of 13 new TDases. The X-ray crystal structures of two new enzymes from Legionella species were determined, and the ability of anhydrotetracycline to inhibit their tetracycline-inactivating activity was confirmed. Using the MSA-based approach we identified 31 amino acid positions 100% conserved across all known TDase sequences. The roles of these positions were analyzed by alanine-scanning mutagenesis in two TDases, to study the impact on cell and in vitro activity, structure, and stability. These results expand the diversity of TDase sequences and provide valuable insights into the roles of important residues in TDases, and flavin monooxygenases more broadly.

60 APPLIED LIFE SCIENCES↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Improved deep learning prediction of antigen–antibody interactions

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.

Science & Technology - Other Topics↗

TopPICR: A Companion R Package for Top-Down Proteomics Data Analysis

Top-down proteomics is the analysis of proteins in their intact form without proteolysis, thus preserving valuable information about post-translational modifications, isoforms, and proteolytic processing. However, it is still a developing field due to limitations in the instrumentation, difficulties with interpretation of complex mass spectra, and a lack of well-established quantification approaches. TopPIC is one of the popular tools for proteoform identification. Here we extended its capabilities into label-free proteoform quantification by developing a companion R package (TopPICR). Key steps in the TopPICR pipeline include filtering identifications, inferring a minimal set of protein accessions explaining the observed sequences, aligning retention times, recalibrating measured masses, clustering features across datasets, and finally compiling feature intensities using the match-between-runs approach. The output of the pipeline is an MSnSet object which makes downstream data analysis seamlessly compatible with packages from the Bioconductor project. It also provides the capability for visualizing proteoforms within the context of the parent protein sequence. The functionality of TopPICR is demonstrated on top-down LC-MS/MS datasets of 10 human-in-mouse xenografts of luminal and basal breast tumor samples.

59 BASIC BIOLOGICAL SCIENCES↗

Instruction Roofline: An insightful visual performance model for GPUs

The Roofline performance model provides an intuitive approach to identify performance bottlenecks and guide performance optimization. However, the classic FLOP-centric approach is inappropriate for the emerging applications that perform more integer operations than floating point operations. In this article, we reintroduce our Instruction Roofline Model on NVIDIA GPUs and expand our evaluation of it. The Instruction Roofline incorporates instructions and memory transactions across all memory hierarchies together, and provides more performance insights than the FLOP-oriented Roofline Model, that is, instruction throughput, stride memory access patterns, bank conflicts, and thread predication. We use our Instruction Roofline methodology to analyze eight proxy applications: HPGMG from AMReX, Matrix Transpose benchmarks, ADEPT from MetaHipMer's sequence alignment phase, EXTENSION from MetaHipMer's local assembly phase, CUSP, cuSPARSE, cudaTensorCoreGemm, and cuBLAS. We demonstrate the ability of our methodology to understand various aspects of performance and performance bottlenecks on NVIDIA GPUs and motivate code optimizations.

Ding, N↗

ppdx : Automated modeling of protein–protein interaction descriptors for use with machine learning

This paper describes ppdx, a python workflow tool that combines protein sequence alignment, homology modeling, and structural refinement, to compute a broad array of descriptors for characterizing protein–protein interactions. The descriptors can be used to predict various properties of interest, such as protein–protein binding affinities, or inhibitory concentrations (IC 50 ), using approaches that range from simple regression to more complex machine learning models. The software is highly modular. It supports different protocols for generating structures, and 95 descriptors can be currently computed. More protocols and descriptors can be easily added. The implementation is highly parallel and can fully exploit the available cores in a single workstation, or multiple nodes on a supercomputer, allowing many systems to be analyzed simultaneously. As an illustrative application, ppdx is used to parametrize a model that predicts the IC 50 of a set of antigens and a class of antibodies directed to the influenza hemagglutinin stalk.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving protein tertiary structure prediction by deep learning and distance prediction in CASP14

Abstract Substantial progresses in protein structure prediction have been made by utilizing deep‐learning and residue‐residue distance prediction since CASP13. Inspired by the advances, we improve our CASP14 MULTICOM protein structure prediction system by incorporating three new components: (a) a new deep learning‐based protein inter‐residue distance predictor to improve template‐free (ab initio) tertiary structure prediction, (b) an enhanced template‐based tertiary structure prediction method, and (c) distance‐based model quality assessment methods empowered by deep learning. In the 2020 CASP14 experiment, MULTICOM predictor was ranked seventh out of 146 predictors in tertiary structure prediction and ranked third out of 136 predictors in inter‐domain structure prediction. The results demonstrate that the template‐free modeling based on deep learning and residue‐residue distance prediction can predict the correct topology for almost all template‐based modeling targets and a majority of hard targets (template‐free targets or targets whose templates cannot be recognized), which is a significant improvement over the CASP13 MULTICOM predictor. Moreover, the template‐free modeling performs better than the template‐based modeling on not only hard targets but also the targets that have homologous templates. The performance of the template‐free modeling largely depends on the accuracy of distance prediction closely related to the quality of multiple sequence alignments. The structural model quality assessment works well on targets for which enough good models can be predicted, but it may perform poorly when only a few good models are predicted for a hard target and the distribution of model quality scores is highly skewed. MULTICOM is available at https://github.com/jianlin-cheng/MULTICOM_Human_CASP14/tree/CASP14_DeepRank3 and https://github.com/multicom-toolbox/multicom/tree/multicom_v2.0 .

59 BASIC BIOLOGICAL SCIENCES↗

Comparison of PsbQ and Psb27 in photosystem II provides insight into their roles

Photosystem II (PSII) catalyzes the oxidation of water at its active site that harbors a high-valent inorganic Mn 4 CaO x cluster called the oxygen-evolving complex (OEC). Extrinsic subunits generally serve to protect the OEC from reductants and stabilize the structure, but diversity in the extrinsic subunits exists between phototrophs. Recent cryo-electron microscopy experiments have provided new molecular structures of PSII with varied extrinsic subunits. We focus on the extrinsic subunit PsbQ, that binds to the mature PSII complex, and on Psb27, an extrinsic subunit involved in PSII biogenesis. PsbQ and Psb27 share a similar binding site and have a four-helix bundle tertiary structure, suggesting they are related. Here, we use sequence alignments, structural analyses, and binding simulations to compare PsbQ and Psb27 from different organisms. We find no evidence that PsbQ and Psb27 are related despite their similar structures and binding sites. Evolutionary divergence within PsbQ homologs from different lineages is high, probably due to their interactions with other extrinsic subunits that themselves exhibit vast diversity between lineages. This may result in functional variation as exemplified by large differences in their calculated binding energies. Psb27 homologs generally exhibit less divergence, which may be due to stronger evolutionary selection for certain residues that maintain its function during PSII biogenesis which is consistent with their more similar calculated binding energies between organisms. Previous experimental inconsistencies, low confidence binding simulations, and recent structural data suggest that Psb27 is likely to exhibit flexibility that may be an important characteristic of its activity. Furthermore, the analysis provides insight into the functions and evolution of PsbQ and Psb27, and an unusual example of proteins with similar tertiary structures and binding sites that probably serve different roles.

59 BASIC BIOLOGICAL SCIENCES↗

A guayule C-repeat binding factor is highly activated in guayule under freezing temperature and enhances freezing tolerance when expressed in Arabidopsis thaliana

Natural Rubber (NR)-producing guayule (Parthenium argentatum Gray) has been developed as a new crop to diversify NR production. Guayule NR is mainly synthesized in its stem and is upregulated by cold temperatures. A guayule C-repeat binding factor 4 (PaCBF4) was highly expressed in cold-treated stem tissue, coinciding with active rubber biosynthesis and accumulation. Sequence alignments of PaCBF4 with other CBFs indicated that PaCBF4 contains DNA-binding domains responsible for regulating cold-regulated (COR) gene expression. Spatial gene expression profiling of PaCBF4 revealed that stems had the highest expression level among different organs examined. We further confirmed the function of PaCBF4 as regulator of cold-signaling processes by expressing it in the model plant Arabidopsis under a constitutive ubiquitin promoter from potato. Further, the resulting transgenic Arabidopsis lines expressing PaCBF4 turned on expression of a set of Arabidopsis COR genes under both room (24ºC) and cold (4ºC) temperatures, in contrast to the wild-type Arabidopsis that expressed these COR genes solely upon cold treatment. Furthermore, the transgenic plants displayed enhanced freezing tolerance at -5ºC, exhibiting a survival rate of 88–98% compared with 0% survival rate of wild-type plants. Our results suggest that PaCBF4 is a functional member of the guayule CBF gene family and plays a significant role in cold and freeze tolerance. Interestingly, overexpressing PaCBF4 in Arabidopsis did not affect the normal phenotype of the plant during vegetative and inflorescence growth, but the gene led to more undeveloped siliques after flowering.

60 APPLIED LIFE SCIENCES↗

Exploring drought-responsive crucial genes in Sorghum

Drought severely affects global food production. Sorghum is a typical drought-resistant model crop. Based on RNA-seq data for Sorghum with multiple time points and the gray correlation coefficient, this paper firstly selects candidate genes via mean variance test and constructs weighted gene differential co-expression networks (WGDCNs); then, based on guilt-by-rewiring principle, the WGDCNs and the hidden Markov random field model, drought-responsive crucial genes are identified for five developmental stages respectively. Enrichment and sequence alignment analysis reveal that the screened genes may play critical functional roles in drought responsiveness. A multilayer differential co-expression network for the screened genes reveals that Sorghum is very sensitive to pre-flowering drought. Furthermore, a crucial gene regulatory module is established, which regulates drought responsiveness via plant hormone signal transduction, MAPK cascades, and transcriptional regulations. The proposed method can well excavate crucial genes through RNA-seq data, which have implications in breeding of new varieties with improved drought tolerance.

60 APPLIED LIFE SCIENCES↗

Engineering and Application of a Thermostable MHETase for PET Depolymerization

Enzymatic hydrolysis of poly(ethylene terephthalate) (PET) releases mono(2-hydroxyethyl) terephthalate (MHET) as a major product, the accumulation of which can prolong reactor residence times and complicate downstream monomer separations. The use of a MHETase enzyme can enable MHET hydrolysis to the monomers, terephthalic acid and ethylene glycol, but industrial PETases typically operate at thermophilic temperatures and the well-known MHETase from Ideonella sakaiensis is a mesophilic enzyme, thus warranting the development of thermophilic MHETases. Here, we characterize thermostable MHET-active enzymes from a natural diversity screen by applying a hidden Markov model based on the previously reported, archaeal ferulic acid esterase, PET46. We identified enzymes with higher thermostability than PET46 and quantified their MHETase activity in reactions at 70 °C. The crystal structure of MHT077, the homologue with the highest MHETase activity and an apparent melting temperature (T m,app ) of 94.6 °C, informed site saturation mutagenesis in the active site and lid-domain interface. MHT077 exhibited a ∼100-fold slower unfolding rate at 65 °C than PET46, indicating substantially greater kinetic stability. In parallel, we applied evolution-informed design, a probabilistic model that leverages coevolutionary patterns in large multiple sequence alignments, to improve the activity and thermostability of five ferulic acid esterases. One design, EV-MHT043–5 was identified with a comparable thermostability (T m,app = 96.1 °C) and a 3-fold improvement in its MHETase activity relative to the wildtype enzyme, MHT043. Combination variants of beneficial mutations were screened and afforded a variant, MHT077 LFK , which reduced MHET accumulation in bioreactor experiments with postconsumer PET waste. Overall, this study expands the known MHET-hydrolyzing protein scaffolds available for enzymatic PET recycling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AF2Complex predicts direct physical interactions in multimeric proteins with deep learning

Abstract Accurate descriptions of protein-protein interactions are essential for understanding biological systems. Remarkably accurate atomic structures have been recently computed for individual proteins by AlphaFold2 (AF2). Here, we demonstrate that the same neural network models from AF2 developed for single protein sequences can be adapted to predict the structures of multimeric protein complexes without retraining. In contrast to common approaches, our method, AF2Complex, does not require paired multiple sequence alignments. It achieves higher accuracy than some complex protein-protein docking strategies and provides a significant improvement over AF-Multimer, a development of AlphaFold for multimeric proteins. Moreover, we introduce metrics for predicting direct protein-protein interactions between arbitrary protein pairs and validate AF2Complex on some challenging benchmark sets and the E. coli proteome. Lastly, using the cytochrome c biogenesis system I as an example, we present high-confidence models of three sought-after assemblies formed by eight members of this system.

59 BASIC BIOLOGICAL SCIENCES↗

Crystal structure of the CoV-Y domain of SARS-CoV-2 nonstructural protein 3

Abstract Replication of the coronavirus genome starts with the formation of viral RNA-containing double-membrane vesicles (DMV) following viral entry into the host cell. The multi-domain nonstructural protein 3 (nsp3) is the largest protein encoded by the known coronavirus genome and serves as a central component of the viral replication and transcription machinery. Previous studies demonstrated that the highly-conserved C-terminal region of nsp3 is essential for subcellular membrane rearrangement, yet the underlying mechanisms remain elusive. Here we report the crystal structure of the CoV-Y domain, the most C-terminal domain of the SARS-CoV-2 nsp3, at 2.4 Å-resolution. CoV-Y adopts a previously uncharacterized V-shaped fold featuring three distinct subdomains. Sequence alignment and structure prediction suggest that this fold is likely shared by the CoV-Y domains from closely related nsp3 homologs. NMR-based fragment screening combined with molecular docking identifies surface cavities in CoV-Y for interaction with potential ligands and other nsps. These studies provide the first structural view on a complete nsp3 CoV-Y domain, and the molecular framework for understanding the architecture, assembly and function of the nsp3 C-terminal domains in coronavirus replication. Our work illuminates nsp3 as a potential target for therapeutic interventions to aid in the on-going battle against the COVID-19 pandemic and diseases caused by other coronaviruses.

36 MATERIALS SCIENCE↗

Improving AlphaFold2-based protein tertiary structure prediction with MULTICOM in CASP15

Since the 14th Critical Assessment of Techniques for Protein Structure Prediction (CASP14), AlphaFold2 has become the standard method for protein tertiary structure prediction. One remaining challenge is to further improve its prediction. We developed a new version of the MULTICOM system to sample diverse multiple sequence alignments (MSAs) and structural templates to improve the input for AlphaFold2 to generate structural models. The models are then ranked by both the pairwise model similarity and AlphaFold2 self-reported model quality score. The top ranked models are refined by a novel structure alignment-based refinement method powered by Foldseek. Moreover, for a monomer target that is a subunit of a protein assembly (complex), MULTICOM integrates tertiary and quaternary structure predictions to account for tertiary structural changes induced by protein-protein interaction. The system participated in the tertiary structure prediction in 2022 CASP15 experiment. Our server predictor MULTICOM_refine ranked 3rd among 47 CASP15 server predictors and our human predictor MULTICOM ranked 7th among all 132 human and server predictors. The average GDT-TS score and TM-score of the first structural models that MULTICOM_refine predicted for 94 CASP15 domains are ~0.80 and ~0.92, 9.6% and 8.2% higher than ~0.73 and 0.85 of the standard AlphaFold2 predictor respectively.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗