Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Bioinformatics and 3D Structural Analysis of the Coronavirus Main Protease Active Site Diversity

Coronaviruses (Coronaviridae) such as SARS‐CoV‐2 (severe acute respiratory syndrome coronavirus) and MERS‐CoV (Middle East respiratory syndrome coronavirus) have been the source of recent outbreaks and global health concerns. While vaccines have been essential for controlling the SARS‐CoV‐2 (COVID‐19) pandemic, it is uncertain whether they will be effective against future coronavirus strains. Therefore, identification or design of a broad‐spectrum drug that targets highly conserved regions of the main protease of multiple coronavirus strains is essential in the long term. As part of a virtual summer research experience with the RCSB PDB, bioinformatics tools were employed to predict and construct 3D models of the coronavirus main protease (MPro) using SARS‐CoV‐2 as the template, with a focus on mutational trends and active sites. This study focused on the active sites of MPro, a cysteine protease essential for viral assembly and replication. Sequence alignments and structure modeling of MPro structures has identified conserved regions across multiple coronavirus strains. Inhibition of MPro halts coronavirus replication, making it an ideal drug target, and studies of MPro may foster and accelerate the discovery of high affinity broad‐spectrum drugs.

Wu Wu, Amy↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗

A fast comparative genome browser for diverse bacteria and archaea

Genome sequencing has revealed an incredible diversity of bacteria and archaea, but there are no fast and convenient tools for browsing across these genomes. It is cumbersome to view the prevalence of homologs for a protein of interest, or the gene neighborhoods of those homologs, across the diversity of the prokaryotes. We developed a web-based tool, fast . genomics , that uses two strategies to support fast browsing across the diversity of prokaryotes. First, the database of genomes is split up. The main database contains one representative from each of the 6,377 genera that have a high-quality genome, and additional databases for each taxonomic order contain up to 10 representatives of each species. Second, homologs of proteins of interest are identified quickly by using accelerated searches, usually in a few seconds. Once homologs are identified, fast . genomics can quickly show their prevalence across taxa, view their neighboring genes, or compare the prevalence of two different proteins. Fast . genomics is available at https://fast.genomics.lbl.gov .

59 BASIC BIOLOGICAL SCIENCES↗

First report of Seville root-knot nematode, Meloidogyne hispanica (Nematoda: Meloidogynidae) in the USA and North America

A high number of second stage juveniles of the root-knot nematode were recovered from soil samples collected from a corn field, located in Pickens County, South Carolina, USA in 2019. Extracted nematodes were examined morphologically and molecularly for species identification which indicated that the specimens of root knot juveniles were Meloidogyne hispanica. The morphological examination and morphometric details from second-stage juveniles were consistent with the original description and redescriptions of this species. The ITS rRNA, D2-D3 expansion segments of 28S rRNA, intergenic COII-16S region, nad5 and COI gene sequences were obtained from the South Carolina population of M. hispanica. Phylogenetic analysis of the intergenic COII-16S region of mtDNA gene sequence alignment using statistical parsimony showed that the South Carolina population clustered with Meloidogyne hispanica from Portugal and Australia. To our best knowledge, this finding represents the first report of Meloidogyne hispanica in the USA and North America.

59 BASIC BIOLOGICAL SCIENCES↗

pH Homeostasis and Sodium Ion Pumping by Multiple Resistance and pH Antiporters in Pyrococcus furiosus

Multiple Resistance and pH (Mrp) antiporters are seven-subunit complexes that couple transport of ions across the membrane in response to a proton motive force (PMF) and have various physiological roles, including sodium ion sensing and pH homeostasis. The hyperthermophilic archaeon Pyrococcus furiosus contains three copies of Mrp encoding genes in its genome. Two are found as integral components of two respiratory complexes, membrane bound hydrogenase (MBH) and the membrane bound sulfane sulfur reductase (MBS) that couple redox activity to sodium translocation, while the third copy is a stand-alone Mrp. Sequence alignments show that this Mrp does not contain an energy-input (PMF) module but contains all other predicted functional Mrp domains. The P. furiosus Mrp deletion strain exhibits no significant changes in optimal pH or sodium ion concentration for growth but is more sensitive to medium acidification during growth. Cell suspension hydrogen gas production assays using the deletion strain show that this Mrp uses sodium as the coupling ion. Mrp likely maintains cytoplasmic pH by exchanging protons inside the cell for extracellular sodium ions. Deletion of the MBH sodium-translocating module demonstrates that hydrogen gas production is uncoupled from ion pumping and provides insights into the evolution of this Mrp-containing respiratory complex.

59 BASIC BIOLOGICAL SCIENCES↗

DeepComplex: A Web Server of Predicting Protein Complex Structures by Deep Learning Inter-chain Contact Prediction and Distance-Based Modelling

Proteins interact to form complexes. Predicting the quaternary structure of protein complexes is useful for protein function analysis, protein engineering, and drug design. However, few user-friendly tools leveraging the latest deep learning technology for inter-chain contact prediction and the distance-based modelling to predict protein quaternary structures are available. To address this gap, we develop DeepComplex, a web server for predicting structures of dimeric protein complexes. It uses deep learning to predict inter-chain contacts in a homodimer or heterodimer. The predicted contacts are then used to construct a quaternary structure of the dimer by the distance-based modelling, which can be interactively viewed and analysed. The web server is freely accessible and requires no registration. It can be easily used by providing a job name and an email address along with the tertiary structure for one chain of a homodimer or two chains of a heterodimer. The output webpage provides the multiple sequence alignment, predicted inter-chain residue-residue contact map, and predicted quaternary structure of the dimer.

59 BASIC BIOLOGICAL SCIENCES↗

Zam Is a Redox-Regulated Member of the RNB-Family Required for Optimal Photosynthesis in Cyanobacteria

The zam gene mediating resistance to acetazolamide in cyanobacteria was discovered thirty years ago during a drug tolerance screen. We use phylogenetics to show that Zam proteins are distributed across cyanobacteria and that they form their own unique clade of the ribonuclease II/R (RNB) family. Despite being RNB family members, multiple sequence alignments reveal that Zam proteins lack conservation and exhibit extreme degeneracy in the canonical active site—raising questions about their cellular function(s). Several known phenotypes arise from the deletion of zam, including drug resistance, slower growth, and altered pigmentation. Using room-temperature and low-temperature fluorescence and absorption spectroscopy, we show that deletion of zam results in decreased phycocyanin synthesis rates, altered PSI:PSII ratios, and an increase in coupling between the phycobilisome and PSII. Conserved cysteines within Zam are identified and assayed for function using in vitro and in vivo methods. We show that these cysteines are essential for Zam function, with mutation of either residue to serine causing phenotypes identical to the deletion of Zam. Redox regulation of Zam activity based on the reversible oxidation-reduction of a disulfide bond involving these cysteine residues could provide a mechanism to integrate the ‘central dogma’ with photosynthesis in cyanobacteria.

59 BASIC BIOLOGICAL SCIENCES↗

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

A General Framework to Learn Tertiary Structure for Protein Sequence Characterization

During the past five years, deep-learning algorithms have enabled ground-breaking progress towards the prediction of tertiary structure from a protein sequence. Very recently, we developed SAdLSA, a new computational algorithm for protein sequence comparison via deep-learning of protein structural alignments. SAdLSA shows significant improvement over established sequence alignment methods. In this contribution, we show that SAdLSA provides a general machine-learning framework for structurally characterizing protein sequences. By aligning a protein sequence against itself, SAdLSA generates a fold distogram for the input sequence, including challenging cases whose structural folds were not present in the training set. About 70% of the predicted distograms are statistically significant. Although at present the accuracy of the intra-sequence distogram predicted by SAdLSA self-alignment is not as good as deep-learning algorithms specifically trained for distogram prediction, it is remarkable that the prediction of single protein structures is encoded by an algorithm that learns ensembles of pairwise structural comparisons, without being explicitly trained to recognize individual structural folds. As such, SAdLSA can not only predict protein folds for individual sequences, but also detects subtle, yet significant, structural relationships between multiple protein sequences using the same deep-learning neural network. The former reduces to a special case in this general framework for protein sequence annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence Impedance Measurement of Utility-Scale Wind Turbines and Inverters – Reference Frame, Frequency Coupling, and MIMO/SISO Forms

Sequence impedance responses with or without considering frequency coupling in both MIMO and SISO forms are increasingly used for the stability analysis of three-phase power electronic systems; however, many aspects of sequence impedance measurement are not fully explored. It is not clear if the sequence impedance has a reference frame similar to the dq impedance. If so, the role of the grid voltage angle estimation in aligning the sequence impedance reference frame has not been discussed. Additionally, existing methods for measuring the sequence impedance with frequency coupling are complicated, are not feasible for large wind turbines and inverters, and provide the sequence impedance responses in only either MIMO or SISO form. This paper presents a sequence impedance measurement method that considers the frequency coupling, performs reference frame alignment, demonstrates the impact of the grid voltage angle estimation, and obtains the sequence impedance response in both MIMO and SISO forms. This paper demonstrates the proposed method and practical problems associated with the sequence impedance measurement of utility-scale wind turbines and inverters on a 1.9-MW Type III wind turbine and a 2.2-MVA inverter using an impedance measurement system built around a 7-MW/13.8-kV grid simulator and a 5-MW dynamometer.

17 WIND ENERGY↗

LoVoCCS. II. Weak Lensing Mass Distributions, Red-sequence Galaxy Distributions, and Their Alignment with the Brightest Cluster Galaxy in 58 Nearby X-Ray-luminous Galaxy Clusters

The Local Volume Complete Cluster Survey is an ongoing program to observe nearly a hundred low-redshift X-ray-luminous galaxy clusters (redshifts 0.03 < z < 0.12 and X-ray luminosities in the 0.1–2.4 keV band L X500c > 10 44 erg s −1 ) with the Dark Energy Camera, capturing data in the u, g, r, i, z bands with a 5σ point source depth of approximately 25th–26th AB magnitudes. Here, we map the aperture masses in 58 galaxy cluster fields using weak gravitational lensing. These clusters span a variety of dynamical states, from nearly relaxed to merging systems, and approximately half of them have not been subject to detailed weak lensing analysis before. In each cluster field, we analyze the alignment between the 2D mass distribution described by the aperture mass map, the 2D red-sequence (RS) galaxy distribution, and the brightest cluster galaxy (BCG). We find that the orientations of the BCG and the RS distribution are strongly aligned throughout the interiors of the clusters: the median misalignment angle is 19° within 2 Mpc. We also observe the alignment between the orientations of the RS distribution and the overall cluster mass distribution (by a median difference of 32° within 1 Mpc), although this is constrained by galaxy shape noise and the limitations of our cluster sample size. These types of alignment suggest long-term dynamical evolution within the clusters over cosmic timescales.

79 ASTRONOMY AND ASTROPHYSICS↗

Multiplex detection and identification of viral, bacterial, and protozoan pathogens in human blood and plasma using an expanded high-density resequencing microarray platform

Introduction: Nucleic acid tests for blood donor screening have improved the safety of the blood supply; however, increasing numbers of emerging pathogen tests are burdensome. Multiplex testing platforms are a potential solution. Methods: The Blood Borne Pathogen Resequencing Microarray Expanded (BBP-RMAv.2) can perform multiplex detection and identification of 80 viruses, bacteria and parasites. This study evaluated pathogen detection in human blood or plasma. Samples spiked with selected pathogens, each with one of 6 viruses, 2 bacteria and 5 protozoans were tested on this platform. The nucleic acids were extracted, amplified using multiplexed sets of primers, and hybridized to a microarray. The reported sequences were aligned to a database to identify the pathogen. To directly compare the microarray to an emerging molecular approach, the amplified nucleic acids were also submitted to nanopore next generation sequencing (NGS). Results: The BBP-RMAv.2 detected viral pathogens at a concentration as low as 100 copies/ml and a range of concentrations from 1,000 to 100,000 copies/ml for all the spiked pathogens. Coded specimens were identified correctly demonstrating the effectiveness of the platform. The nanopore sequencing correctly identified most samples and the results of the two platforms were compared. Discussion: These results indicated that the BBP-RMAv.2 could be employed for multiplex detection with potential for use in blood safety or disease diagnosis. The NGS was nearly as effective at identifying pathogens in blood and performed better than BBP-RMAv.2 at identifying pathogen-negative samples.

59 BASIC BIOLOGICAL SCIENCES↗

Classification of bacterial plasmid and chromosome derived sequences using machine learning

Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.

59 BASIC BIOLOGICAL SCIENCES↗

SCOPe: improvements to the structural classification of proteins – extended database to facilitate variant interpretation and machine learning

Abstract The Structural Classification of Proteins—extended (SCOPe, https://scop.berkeley.edu) knowledgebase aims to provide an accurate, detailed, and comprehensive description of the structural and evolutionary relationships amongst the majority of proteins of known structure, along with resources for analyzing the protein structures and their sequences. Structures from the PDB are divided into domains and classified using a combination of manual curation and highly precise automated methods. In the current release of SCOPe, 2.08, we have developed search and display tools for analysis of genetic variants we mapped to structures classified in SCOPe. In order to improve the utility of SCOPe to automated methods such as deep learning classifiers that rely on multiple alignment of sequences of homologous proteins, we have introduced new machine-parseable annotations that indicate aberrant structures as well as domains that are distinguished by a smaller repeat unit. We also classified structures from 74 of the largest Pfam families not previously classified in SCOPe, and we improved our algorithm to remove N- and C-terminal cloning, expression and purification sequences from SCOPe domains. SCOPe 2.08-stable classifies 106 976 PDB entries (about 60% of PDB entries).

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of genomic signatures associated with Variovorax endosphere colonization

This repository contains the analysis code and supporting datasets associated with the study “Genomic signatures in Variovorax enabling colonization of the Populus endosphere.” Beals DG, Carper DL, Hochanadel LH, Jawdy SS, Klingeman DM, Piatkowski BT, Weston DJ, Doktycz MJ, Pelletier DA. 2026. Genomic signatures in Variovorax enabling colonization of the Populus endosphere. mSystems 11:e01605-25. https://doi.org/10.1128/msystems.01605-25 The scripts are organized sequentially (01–07) and document the workflows used for: Sequence-read alignment and feature counting Orthogroup and KEGG Ortholog annotation Count normalization Statistical analysis and aggregation Generation of manuscript figures and tables Repository contents The uncompressed files are the finalized, formatted datasets used to generate the figures and tables reported in the study, including the supplemental CSV files referenced in the manuscript. The accompanying ZIP archive contains the complete codebase and example data_input/ and data_output/ directories illustrating the organization and execution of the analytical workflow. Individual scripts identify the corresponding manuscript analyses and figure panels. Raw sequencing data Raw sequencing reads are available through the NCBI Sequence Read Archive under BioProject accession PRJNA1322484.

Beals, Delaney [ORNL] (ORCID:0000000306274574)↗