Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A guayule C-repeat binding factor is highly activated in guayule under freezing temperature and enhances freezing tolerance when expressed in Arabidopsis thaliana

Natural Rubber (NR)-producing guayule (Parthenium argentatum Gray) has been developed as a new crop to diversify NR production. Guayule NR is mainly synthesized in its stem and is upregulated by cold temperatures. A guayule C-repeat binding factor 4 (PaCBF4) was highly expressed in cold-treated stem tissue, coinciding with active rubber biosynthesis and accumulation. Sequence alignments of PaCBF4 with other CBFs indicated that PaCBF4 contains DNA-binding domains responsible for regulating cold-regulated (COR) gene expression. Spatial gene expression profiling of PaCBF4 revealed that stems had the highest expression level among different organs examined. We further confirmed the function of PaCBF4 as regulator of cold-signaling processes by expressing it in the model plant Arabidopsis under a constitutive ubiquitin promoter from potato. Further, the resulting transgenic Arabidopsis lines expressing PaCBF4 turned on expression of a set of Arabidopsis COR genes under both room (24ºC) and cold (4ºC) temperatures, in contrast to the wild-type Arabidopsis that expressed these COR genes solely upon cold treatment. Furthermore, the transgenic plants displayed enhanced freezing tolerance at -5ºC, exhibiting a survival rate of 88–98% compared with 0% survival rate of wild-type plants. Our results suggest that PaCBF4 is a functional member of the guayule CBF gene family and plays a significant role in cold and freeze tolerance. Interestingly, overexpressing PaCBF4 in Arabidopsis did not affect the normal phenotype of the plant during vegetative and inflorescence growth, but the gene led to more undeveloped siliques after flowering.

60 APPLIED LIFE SCIENCES↗

Exploring drought-responsive crucial genes in Sorghum

Drought severely affects global food production. Sorghum is a typical drought-resistant model crop. Based on RNA-seq data for Sorghum with multiple time points and the gray correlation coefficient, this paper firstly selects candidate genes via mean variance test and constructs weighted gene differential co-expression networks (WGDCNs); then, based on guilt-by-rewiring principle, the WGDCNs and the hidden Markov random field model, drought-responsive crucial genes are identified for five developmental stages respectively. Enrichment and sequence alignment analysis reveal that the screened genes may play critical functional roles in drought responsiveness. A multilayer differential co-expression network for the screened genes reveals that Sorghum is very sensitive to pre-flowering drought. Furthermore, a crucial gene regulatory module is established, which regulates drought responsiveness via plant hormone signal transduction, MAPK cascades, and transcriptional regulations. The proposed method can well excavate crucial genes through RNA-seq data, which have implications in breeding of new varieties with improved drought tolerance.

60 APPLIED LIFE SCIENCES↗

Engineering and Application of a Thermostable MHETase for PET Depolymerization

Enzymatic hydrolysis of poly(ethylene terephthalate) (PET) releases mono(2-hydroxyethyl) terephthalate (MHET) as a major product, the accumulation of which can prolong reactor residence times and complicate downstream monomer separations. The use of a MHETase enzyme can enable MHET hydrolysis to the monomers, terephthalic acid and ethylene glycol, but industrial PETases typically operate at thermophilic temperatures and the well-known MHETase from Ideonella sakaiensis is a mesophilic enzyme, thus warranting the development of thermophilic MHETases. Here, we characterize thermostable MHET-active enzymes from a natural diversity screen by applying a hidden Markov model based on the previously reported, archaeal ferulic acid esterase, PET46. We identified enzymes with higher thermostability than PET46 and quantified their MHETase activity in reactions at 70 °C. The crystal structure of MHT077, the homologue with the highest MHETase activity and an apparent melting temperature (T m,app ) of 94.6 °C, informed site saturation mutagenesis in the active site and lid-domain interface. MHT077 exhibited a ∼100-fold slower unfolding rate at 65 °C than PET46, indicating substantially greater kinetic stability. In parallel, we applied evolution-informed design, a probabilistic model that leverages coevolutionary patterns in large multiple sequence alignments, to improve the activity and thermostability of five ferulic acid esterases. One design, EV-MHT043–5 was identified with a comparable thermostability (T m,app = 96.1 °C) and a 3-fold improvement in its MHETase activity relative to the wildtype enzyme, MHT043. Combination variants of beneficial mutations were screened and afforded a variant, MHT077 LFK , which reduced MHET accumulation in bioreactor experiments with postconsumer PET waste. Overall, this study expands the known MHET-hydrolyzing protein scaffolds available for enzymatic PET recycling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AF2Complex predicts direct physical interactions in multimeric proteins with deep learning

Abstract Accurate descriptions of protein-protein interactions are essential for understanding biological systems. Remarkably accurate atomic structures have been recently computed for individual proteins by AlphaFold2 (AF2). Here, we demonstrate that the same neural network models from AF2 developed for single protein sequences can be adapted to predict the structures of multimeric protein complexes without retraining. In contrast to common approaches, our method, AF2Complex, does not require paired multiple sequence alignments. It achieves higher accuracy than some complex protein-protein docking strategies and provides a significant improvement over AF-Multimer, a development of AlphaFold for multimeric proteins. Moreover, we introduce metrics for predicting direct protein-protein interactions between arbitrary protein pairs and validate AF2Complex on some challenging benchmark sets and the E. coli proteome. Lastly, using the cytochrome c biogenesis system I as an example, we present high-confidence models of three sought-after assemblies formed by eight members of this system.

59 BASIC BIOLOGICAL SCIENCES↗

Crystal structure of the CoV-Y domain of SARS-CoV-2 nonstructural protein 3

Abstract Replication of the coronavirus genome starts with the formation of viral RNA-containing double-membrane vesicles (DMV) following viral entry into the host cell. The multi-domain nonstructural protein 3 (nsp3) is the largest protein encoded by the known coronavirus genome and serves as a central component of the viral replication and transcription machinery. Previous studies demonstrated that the highly-conserved C-terminal region of nsp3 is essential for subcellular membrane rearrangement, yet the underlying mechanisms remain elusive. Here we report the crystal structure of the CoV-Y domain, the most C-terminal domain of the SARS-CoV-2 nsp3, at 2.4 Å-resolution. CoV-Y adopts a previously uncharacterized V-shaped fold featuring three distinct subdomains. Sequence alignment and structure prediction suggest that this fold is likely shared by the CoV-Y domains from closely related nsp3 homologs. NMR-based fragment screening combined with molecular docking identifies surface cavities in CoV-Y for interaction with potential ligands and other nsps. These studies provide the first structural view on a complete nsp3 CoV-Y domain, and the molecular framework for understanding the architecture, assembly and function of the nsp3 C-terminal domains in coronavirus replication. Our work illuminates nsp3 as a potential target for therapeutic interventions to aid in the on-going battle against the COVID-19 pandemic and diseases caused by other coronaviruses.

36 MATERIALS SCIENCE↗

Improving AlphaFold2-based protein tertiary structure prediction with MULTICOM in CASP15

Since the 14th Critical Assessment of Techniques for Protein Structure Prediction (CASP14), AlphaFold2 has become the standard method for protein tertiary structure prediction. One remaining challenge is to further improve its prediction. We developed a new version of the MULTICOM system to sample diverse multiple sequence alignments (MSAs) and structural templates to improve the input for AlphaFold2 to generate structural models. The models are then ranked by both the pairwise model similarity and AlphaFold2 self-reported model quality score. The top ranked models are refined by a novel structure alignment-based refinement method powered by Foldseek. Moreover, for a monomer target that is a subunit of a protein assembly (complex), MULTICOM integrates tertiary and quaternary structure predictions to account for tertiary structural changes induced by protein-protein interaction. The system participated in the tertiary structure prediction in 2022 CASP15 experiment. Our server predictor MULTICOM_refine ranked 3rd among 47 CASP15 server predictors and our human predictor MULTICOM ranked 7th among all 132 human and server predictors. The average GDT-TS score and TM-score of the first structural models that MULTICOM_refine predicted for 94 CASP15 domains are ~0.80 and ~0.92, 9.6% and 8.2% higher than ~0.73 and 0.85 of the standard AlphaFold2 predictor respectively.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Bioinformatics and 3D Structural Analysis of the Coronavirus Main Protease Active Site Diversity

Coronaviruses (Coronaviridae) such as SARS‐CoV‐2 (severe acute respiratory syndrome coronavirus) and MERS‐CoV (Middle East respiratory syndrome coronavirus) have been the source of recent outbreaks and global health concerns. While vaccines have been essential for controlling the SARS‐CoV‐2 (COVID‐19) pandemic, it is uncertain whether they will be effective against future coronavirus strains. Therefore, identification or design of a broad‐spectrum drug that targets highly conserved regions of the main protease of multiple coronavirus strains is essential in the long term. As part of a virtual summer research experience with the RCSB PDB, bioinformatics tools were employed to predict and construct 3D models of the coronavirus main protease (MPro) using SARS‐CoV‐2 as the template, with a focus on mutational trends and active sites. This study focused on the active sites of MPro, a cysteine protease essential for viral assembly and replication. Sequence alignments and structure modeling of MPro structures has identified conserved regions across multiple coronavirus strains. Inhibition of MPro halts coronavirus replication, making it an ideal drug target, and studies of MPro may foster and accelerate the discovery of high affinity broad‐spectrum drugs.

Wu Wu, Amy↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗

A fast comparative genome browser for diverse bacteria and archaea

Genome sequencing has revealed an incredible diversity of bacteria and archaea, but there are no fast and convenient tools for browsing across these genomes. It is cumbersome to view the prevalence of homologs for a protein of interest, or the gene neighborhoods of those homologs, across the diversity of the prokaryotes. We developed a web-based tool, fast . genomics , that uses two strategies to support fast browsing across the diversity of prokaryotes. First, the database of genomes is split up. The main database contains one representative from each of the 6,377 genera that have a high-quality genome, and additional databases for each taxonomic order contain up to 10 representatives of each species. Second, homologs of proteins of interest are identified quickly by using accelerated searches, usually in a few seconds. Once homologs are identified, fast . genomics can quickly show their prevalence across taxa, view their neighboring genes, or compare the prevalence of two different proteins. Fast . genomics is available at https://fast.genomics.lbl.gov .

59 BASIC BIOLOGICAL SCIENCES↗

First report of Seville root-knot nematode, Meloidogyne hispanica (Nematoda: Meloidogynidae) in the USA and North America

A high number of second stage juveniles of the root-knot nematode were recovered from soil samples collected from a corn field, located in Pickens County, South Carolina, USA in 2019. Extracted nematodes were examined morphologically and molecularly for species identification which indicated that the specimens of root knot juveniles were Meloidogyne hispanica. The morphological examination and morphometric details from second-stage juveniles were consistent with the original description and redescriptions of this species. The ITS rRNA, D2-D3 expansion segments of 28S rRNA, intergenic COII-16S region, nad5 and COI gene sequences were obtained from the South Carolina population of M. hispanica. Phylogenetic analysis of the intergenic COII-16S region of mtDNA gene sequence alignment using statistical parsimony showed that the South Carolina population clustered with Meloidogyne hispanica from Portugal and Australia. To our best knowledge, this finding represents the first report of Meloidogyne hispanica in the USA and North America.

59 BASIC BIOLOGICAL SCIENCES↗

pH Homeostasis and Sodium Ion Pumping by Multiple Resistance and pH Antiporters in Pyrococcus furiosus

Multiple Resistance and pH (Mrp) antiporters are seven-subunit complexes that couple transport of ions across the membrane in response to a proton motive force (PMF) and have various physiological roles, including sodium ion sensing and pH homeostasis. The hyperthermophilic archaeon Pyrococcus furiosus contains three copies of Mrp encoding genes in its genome. Two are found as integral components of two respiratory complexes, membrane bound hydrogenase (MBH) and the membrane bound sulfane sulfur reductase (MBS) that couple redox activity to sodium translocation, while the third copy is a stand-alone Mrp. Sequence alignments show that this Mrp does not contain an energy-input (PMF) module but contains all other predicted functional Mrp domains. The P. furiosus Mrp deletion strain exhibits no significant changes in optimal pH or sodium ion concentration for growth but is more sensitive to medium acidification during growth. Cell suspension hydrogen gas production assays using the deletion strain show that this Mrp uses sodium as the coupling ion. Mrp likely maintains cytoplasmic pH by exchanging protons inside the cell for extracellular sodium ions. Deletion of the MBH sodium-translocating module demonstrates that hydrogen gas production is uncoupled from ion pumping and provides insights into the evolution of this Mrp-containing respiratory complex.

59 BASIC BIOLOGICAL SCIENCES↗

DeepComplex: A Web Server of Predicting Protein Complex Structures by Deep Learning Inter-chain Contact Prediction and Distance-Based Modelling

Proteins interact to form complexes. Predicting the quaternary structure of protein complexes is useful for protein function analysis, protein engineering, and drug design. However, few user-friendly tools leveraging the latest deep learning technology for inter-chain contact prediction and the distance-based modelling to predict protein quaternary structures are available. To address this gap, we develop DeepComplex, a web server for predicting structures of dimeric protein complexes. It uses deep learning to predict inter-chain contacts in a homodimer or heterodimer. The predicted contacts are then used to construct a quaternary structure of the dimer by the distance-based modelling, which can be interactively viewed and analysed. The web server is freely accessible and requires no registration. It can be easily used by providing a job name and an email address along with the tertiary structure for one chain of a homodimer or two chains of a heterodimer. The output webpage provides the multiple sequence alignment, predicted inter-chain residue-residue contact map, and predicted quaternary structure of the dimer.

59 BASIC BIOLOGICAL SCIENCES↗

Zam Is a Redox-Regulated Member of the RNB-Family Required for Optimal Photosynthesis in Cyanobacteria

The zam gene mediating resistance to acetazolamide in cyanobacteria was discovered thirty years ago during a drug tolerance screen. We use phylogenetics to show that Zam proteins are distributed across cyanobacteria and that they form their own unique clade of the ribonuclease II/R (RNB) family. Despite being RNB family members, multiple sequence alignments reveal that Zam proteins lack conservation and exhibit extreme degeneracy in the canonical active site—raising questions about their cellular function(s). Several known phenotypes arise from the deletion of zam, including drug resistance, slower growth, and altered pigmentation. Using room-temperature and low-temperature fluorescence and absorption spectroscopy, we show that deletion of zam results in decreased phycocyanin synthesis rates, altered PSI:PSII ratios, and an increase in coupling between the phycobilisome and PSII. Conserved cysteines within Zam are identified and assayed for function using in vitro and in vivo methods. We show that these cysteines are essential for Zam function, with mutation of either residue to serine causing phenotypes identical to the deletion of Zam. Redox regulation of Zam activity based on the reversible oxidation-reduction of a disulfide bond involving these cysteine residues could provide a mechanism to integrate the ‘central dogma’ with photosynthesis in cyanobacteria.

59 BASIC BIOLOGICAL SCIENCES↗

Two Strategies for Microbial Production of an Industrial Enzyme-Alpha-Amylase

Extremophiles are microorganisms that thrive in, from an anthropocentric view, extreme environments including hot springs, soda lakes and arctic water. This ability of survival at extreme conditions has rendered extremophiles to be of interest in astrobiology, evolutionary biology as well as in industrial applications. Of particular interest to the biotechnology industry are the biological catalysts of the extremophiles, the extremozymes, whose unique stabilities at extreme conditions make them potential sources of novel enzymes in industrial applications. There are two major approaches to microbial enzyme production. This entails enzyme isolation directly from the natural host or creating a recombinant expression system whereby the targeted enzyme can be overexpressed in a mesophilic host. We are employing both methods in the effort to produce alpha-amylases from a hyperthermophilic archaeon (Thermococcus) isolated from a hydrothermal vent in the Atlantic Ocean, as well as from alkaliphilic bacteria (Bacillus) isolated from a soda lake in Tanzania. Alpha-amylases catalyze the hydrolysis of internal alpha-1,4-glycosidic linkages in starch to produce smaller sugars. Thermostable alpha-amylases are used in the liquefaction of starch for production of fructose and glucose syrups, whereas alpha-amylases stable at high pH have potential as detergent additives. The alpha-amylase encoding gene from Thermococcus was PCR amplified using carefully designed primers and analyzed using bioinformatics tools such as BLAST and Multiple Sequence Alignment for cloning and expression in E.coli. Four strains of Bacillus were grown in alkaline starch-enriched medium of which the culture supernatant was used as enzyme source. Amylolytic activity was detected using the starch-iodine method.

Bernhardsdotter, Eva C. M. J.↗