Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

Non-NASA Center↗

EMMA: a new method for computing multiple sequence alignments given a constraint subset alignment

Abstract Background Adding sequences into an existing (possibly user-provided) alignment has multiple applications, including updating a large alignment with new data, adding sequences into a constraint alignment constructed using biological knowledge, or computing alignments in the presence of sequence length heterogeneity. Although this is a natural problem, only a few tools have been developed to use this information with high fidelity. Results We present EMMA (Extending Multiple alignments using MAFFT--add) for the problem of adding a set of unaligned sequences into a multiple sequence alignment (i.e., a constraint alignment). EMMA builds on MAFFT--add, which is also designed to add sequences into a given constraint alignment. EMMA improves on MAFFT--add methods by using a divide-and-conquer framework to scale its most accurate version, MAFFT-linsi--add, to constraint alignments with many sequences. We show that EMMA has an accuracy advantage over other techniques for adding sequences into alignments under many realistic conditions and can scale to large datasets with high accuracy (hundreds of thousands of sequences). EMMA is available at https://github.com/c5shen/EMMA . Conclusions EMMA is a new tool that provides high accuracy and scalability for adding sequences into an existing alignment.

Shen, Chengze↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

New alignment method for remote protein sequences by the direct use of pairwise sequence correlations and substitutions

Understanding protein sequences and how they relate to the functions of proteins is extremely important. One of the most basic operations in bioinformatics is sequence alignment and usually the first things learned from these are which positions are the most conserved and often these are critical parts of the structure, such as enzyme active site residues. In addition, the contact pairs in a protein usually correspond closely to the correlations between residue positions in the multiple sequence alignment, and these usually change in a systematic and coordinated way, if one position changes then the other member of the pair also changes to compensate. In the present work, these correlated pairs are taken as anchor points for a new type of sequence alignment. The main advantage of the method here is its combining the remote homolog detection from our method PROST with pairwise sequence substitutions in the rigorous method from Kleinjung et al. We show a few examples of some resulting sequence alignments, and how they can lead to improvements in alignments for function, even for a disordered protein.

Jia, Kejue↗

Enhancing alphafold-multimer-based protein complex structure prediction with MULTICOM in CASP15

To enhance the AlphaFold-Multimer-based protein complex structure prediction, we developed a quaternary structure prediction system (MULTICOM) to improve the input fed to AlphaFold-Multimer and evaluate and refine its outputs. MULTICOM samples diverse multiple sequence alignments (MSAs) and templates for AlphaFold-Multimer to generate structural predictions by using both traditional sequence alignments and Foldseek-based structure alignments, ranks structural predictions through multiple complementary metrics, and refines the structural predictions via a Foldseek structure alignment-based refinement method. The MULTICOM system with different implementations was blindly tested in the assembly structure prediction in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 as both server and human predictors. MULTICOM_qa ranked 3 rd among 26 CASP15 server predictors and MULTICOM_human ranked 7 th among 87 CASP15 server and human predictors. The average TM-score of the first predictions submitted by MULTICOM_qa for CASP15 assembly targets is ~0.76, 5.3% higher than ~0.72 of the standard AlphaFold-Multimer. The average TM-score of the best of top 5 predictions submitted by MULTICOM_qa is ~0.80, about 8% higher than ~0.74 of the standard AlphaFold-Multimer. Moreover, the Foldseek Structure Alignment-based Multimer structure Generation (FSAMG) method outperforms the widely used sequence alignment-based multimer structure generation.

59 BASIC BIOLOGICAL SCIENCES↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Hypermut 3: identifying specific mutational patterns in a defined nucleotide context that allows multistate characters

Abstract Motivation The detection of APOBEC3F- and APOBEC3G-induced mutations in virus sequences is useful for identifying hypermutated sequences. These sequences are not representative of viral evolution and can therefore alter the results of downstream sequence analyses if included. We previously published the software Hypermut, which detects hypermutation events in sequences relative to a reference. Two versions of this method are available as a webtool. Neither of these methods consider multistate characters or gaps in the sequence alignment. Results Here, we present an updated, user-friendly web and command-line version of Hypermut with functionality to handle multistate characters and gaps in the sequence alignment. This tool allows for straightforward integration of hypermutation detection into sequence analysis pipelines. As with the previous tool, while the main purpose is to identify G to A hypermutation events, any mutational pattern and context can be specified. Availability and implementation Hypermut 3 is written in Python 3. It is available as a command-line tool at https://github.com/MolEvolEpid/hypermut3 and as a webtool at https://www.hiv.lanl.gov/content/sequence/HYPERMUT/hypermutv3.html.

59 BASIC BIOLOGICAL SCIENCES↗

Size and Structure of the Sequence Space of Repeat Proteins

The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family— the total number of sequences in that family—can be estimated using models of maximum entropy trained on multiple sequence alignments of naturally occurring amino acid sequences. We analyzed and calculated the size of three abundant repeat proteins families, whose members are large proteins made of many repetitions of conserved portions of *30 amino acids. While amino acid conservation at each position of the alignment explains most of the reduction of diversity relative to completely random sequences, we found that correlations between amino acid usage at different positions significantly impact that diversity. We quantified the impact of different types of correlations, functional and evolutionary, on sequence diversity. Analysis of the detailed structure of the coding space of the families revealed a rugged landscape, with many local energy minima of varying sizes with a hierarchical structure, reminiscent of frustrated energy landscapes of spin glass in physics. This clustered structure indicates a multiplicity of subtypes within each family and suggests new strategies for protein design.

Jacopo Marchi↗

Genomic and morphological characterization of Knufia obscura isolated from the Mars 2020 spacecraft assembly facility

Members of the family Trichomeriaceae, belonging to the Chaetothyriales order and the Ascomycota phylum, are known for their capability to inhabit hostile environments characterized by extreme temperatures, oligotrophic conditions, drought, or presence of toxic compounds. The genus Knufia encompasses many polyextremophilic species. In this report, the genomic and morphological features of the strain FJI-L2-BK-P2 presented, which was isolated from the Mars 2020 mission spacecraft assembly facility located at the Jet Propulsion Laboratory in Pasadena, California. The identification is based on sequence alignment for marker genes, multi-locus sequence analysis, and whole genome sequence phylogeny. The morphological features were studied using a diverse range of microscopic techniques (bright field, phase contrast, differential interference contrast and scanning electron microscopy). The phylogenetic marker genes of the strain FJI-L2-BK-P2 exhibited highest similarities with type strain of Knufia obscura (CBS 148926 T ) that was isolated from the gas tank of a car in Italy. To validate the species identity, whole genomes of both strains (FJI-L2-BK-P2 and CBS 148926 T ) were sequenced, annotated, and strain FJI-L2-BK-P2 was confirmed as K. obscura. The morphological analysis and description of the genomic characteristics of K. obscura FJI-L2-BK-P2 may contribute to refining the taxonomy of Knufia species. Key morphological features are reported in this K. obscura strain, resembling microsclerotia and chlamydospore-like propagules. These features known to be characteristic features in black fungi which could potentially facilitate their adaptation to harsh environments.

59 BASIC BIOLOGICAL SCIENCES↗

Inventory of Composable Elements (ICE) v6.0.0

The Inventory of Composable Elements (ICE) is an open source registry software platform for managing information about biological parts. It is capable of recording information about plasmids, microbial host strains and seeds, as well as DNA parts. Includes features such as DNA sequence visualization, editing and annotation, auto-aligning sequencing trace files against reference templates, SBOL XML/RDF support, and web-of-registries functionality. The web of registries functionality provides strong support for distributed interconnected use and enables sharing and transfer of biological parts across various independent ICE instances. ICE adopts modern software development principles, leveraging component-base frameworks, offering a REST API for convenient third-party integration and emphasizing scalability, security, and service integrations for dynamic content availability. The source code is hosted at https://github.com/JBEI/ice. A public instance is available at public-registry.jbei.org, where users can try out features, upload parts or simply use it for their projects.

Plahar, Hector↗

Cloning and characterization of ftsZ and pyrF from the archaeon Thermoplasma acidophilum

To characterize cytoskeletal components of archaea, the ftsZ gene from Thermoplasma acidophilum was cloned and sequenced. In T. acidophilum ftsZ, which is involved in cell division, was found to be in an operon with the pyrF gene, which encodes orotidine-5'-monophosphate decarboxylase (ODC), an essential enzyme in pyrimidine biosynthesis. Both ftsZ and pyrF from T. acidophilum were expressed in Escherichia coli and formed functional proteins. FtsZ expression in wild-type E. coli resulted in the filamentous phenotype characteristic of ftsZ mutants. T. acidophilum pyrF expression in an E. coli mutant lacking pyrF complemented the mutation and rescued the strain. Sequence alignments of ODCs from archaea, bacteria, and eukarya reveal five conserved regions, two of which have homology to 3-hexulose-6-phosphate synthase (HPS), suggesting a common substrate recognition and binding motif. Copyright 2000 Academic Press.

Bacterial Proteins/chemistry/genetics/metabolism↗

Poplar

SAND2025-00683O Poplar is a software tool that generates a phylogenetic tree from input gene and genome sequences. It integrates established tools to identify genes within genomes, group sequences, construct gene trees, and infer a species tree. Poplar processes nucleotide sequences, identifies similar sequences using Nucleotide BLAST, groups them with DBSCAN, aligns sequences with MAFFT, constructs gene trees with RAxML-NG, and infers a species tree using ASTRAL-Pro3. This pipeline provides a structured approach to phylogenetic analysis, facilitating the study of evolutionary relationships among species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krishnakumar, Raga↗

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Origins of prokaryotes, eukaryotes, mitochondria, and chloroplasts

A computer branching model is used to analyze cellular evolution. Attention is given to certain key amino acids and nucleotide residues (ferredoxin, 5s ribosomal RNA, and c-type cytochromes) because of their commonality over a wide variety of cell types. Each amino acid or nucleotide residue is a sequence in an inherited biological trait; and the branching method is employed to align sequences so that changes reflect substitution of one residue for another. Based on the computer analysis, the symbiotic theory of cellular evolution is considered the most probable. This theory holds that organelles, e.g., mitochondria and chloroplasts invaded larger bodies, e.g., bacteria, and combined functions to form eucaryotic cells.

Schwartz, R. M.↗