Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

SCOPe: improvements to the structural classification of proteins – extended database to facilitate variant interpretation and machine learning

Abstract The Structural Classification of Proteins—extended (SCOPe, https://scop.berkeley.edu) knowledgebase aims to provide an accurate, detailed, and comprehensive description of the structural and evolutionary relationships amongst the majority of proteins of known structure, along with resources for analyzing the protein structures and their sequences. Structures from the PDB are divided into domains and classified using a combination of manual curation and highly precise automated methods. In the current release of SCOPe, 2.08, we have developed search and display tools for analysis of genetic variants we mapped to structures classified in SCOPe. In order to improve the utility of SCOPe to automated methods such as deep learning classifiers that rely on multiple alignment of sequences of homologous proteins, we have introduced new machine-parseable annotations that indicate aberrant structures as well as domains that are distinguished by a smaller repeat unit. We also classified structures from 74 of the largest Pfam families not previously classified in SCOPe, and we improved our algorithm to remove N- and C-terminal cloning, expression and purification sequences from SCOPe domains. SCOPe 2.08-stable classifies 106 976 PDB entries (about 60% of PDB entries).

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of genomic signatures associated with Variovorax endosphere colonization

This repository contains the analysis code and supporting datasets associated with the study “Genomic signatures in Variovorax enabling colonization of the Populus endosphere.” Beals DG, Carper DL, Hochanadel LH, Jawdy SS, Klingeman DM, Piatkowski BT, Weston DJ, Doktycz MJ, Pelletier DA. 2026. Genomic signatures in Variovorax enabling colonization of the Populus endosphere. mSystems 11:e01605-25. https://doi.org/10.1128/msystems.01605-25 The scripts are organized sequentially (01–07) and document the workflows used for: Sequence-read alignment and feature counting Orthogroup and KEGG Ortholog annotation Count normalization Statistical analysis and aggregation Generation of manuscript figures and tables Repository contents The uncompressed files are the finalized, formatted datasets used to generate the figures and tables reported in the study, including the supplemental CSV files referenced in the manuscript. The accompanying ZIP archive contains the complete codebase and example data_input/ and data_output/ directories illustrating the organization and execution of the analytical workflow. Individual scripts identify the corresponding manuscript analyses and figure panels. Raw sequencing data Raw sequencing reads are available through the NCBI Sequence Read Archive under BioProject accession PRJNA1322484.

Beals, Delaney [ORNL] (ORCID:0000000306274574)↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗

Tandem repeats in giant archaeal Borg elements undergo rapid evolution and create new intrinsically disordered regions in proteins

Borgs are huge, linear extrachromosomal elements associated with anaerobic methane-oxidizing archaea. Striking features of Borg genomes are pervasive tandem direct repeat (TR) regions. Here, we present six new Borg genomes and investigate the characteristics of TRs in all ten complete Borg genomes. We find that TR regions are rapidly evolving, recently formed, arise independently, and are virtually absent in host Methanoperedens genomes. Flanking partial repeats and A-enriched character constrain the TR formation mechanism. TRs can be in intergenic regions, where they might serve as regulatory RNAs, or in open reading frames (ORFs). TRs in ORFs are under very strong selective pressure, leading to perfect amino acid TRs (aaTRs) that are commonly intrinsically disordered regions. Proteins with aaTRs are often extracellular or membrane proteins, and functionally similar or homologous proteins often have aaTRs composed of the same amino acids. We propose that Borg aaTR-proteins functionally diversify Methanoperedens and all TRs are crucial for specific Borg–host associations and possibly cospeciation.

59 BASIC BIOLOGICAL SCIENCES↗

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Tracking ebolavirus genomic drift with a resequencing microarray

Filoviruses are emerging pathogens that cause acute fever with high fatality rate and present a global public health threat. During the 2013–2016 Ebola virus outbreak, genome sequencing allowed the study of virus evolution, mutations affecting pathogenicity and infectivity, and tracing the viral spread. In 2018, early sequence identification of the Ebolavirus as EBOV in the Democratic Republic of the Congo supported the use of an Ebola virus vaccine. However, field-deployable sequencing methods are needed to enable a rapid public health response. Resequencing microarrays (RMA) are a targeted method to obtain genomic sequence on clinical specimens rapidly, and sensitively, overcoming the need for extensive bioinformatic analysis. This study presents the design and initial evaluation of an ebolavirus resequencing microarray (Ebolavirus-RMA) system for sequencing the major genomic regions of four Ebolaviruses that cause disease in humans. The design of the Ebolavirus-RMA system is described and evaluated by sequencing repository samples of three Ebolaviruses and two EBOV variants. The ability of the system to identify genetic drift in a replicating virus was achieved by sequencing the ebolavirus glycoprotein gene in a recombinant virus cultured under pressure from a neutralizing antibody. Comparison of the Ebolavirus-RMA results to the Genbank database sequence file with the accession number given for the source RNA and Ebolavirus-RMA results compared to Next Generation Sequence results of the same RNA samples showed up to 99% agreement.

59 BASIC BIOLOGICAL SCIENCES↗

Structure of the T. brucei kinetoplastid RNA editing substrate-binding complex core component, RESC5

Kinetoplastid protists such as Trypanosoma brucei undergo an unusual process of mitochondrial uridine (U) insertion and deletion editing termed kinetoplastid RNA editing (kRNA editing). This extensive form of editing, which is mediated by guide RNAs (gRNAs), can involve the insertion of hundreds of Us and deletion of tens of Us to form a functional mitochondrial mRNA transcript. kRNA editing is catalyzed by the 20 S editosome/RECC. However, gRNA directed, processive editing requires the RNA editing substrate binding complex (RESC), which is comprised of 6 core proteins, RESC1-RESC6. To date there are no structures of RESC proteins or complexes and because RESC proteins show no homology to proteins of known structure, their molecular architecture remains unknown. RESC5 is a key core component in forming the foundation of the RESC complex. To gain insight into the RESC5 protein we performed biochemical and structural studies. We show that RESC5 is monomeric and we report the T . brucei RESC5 crystal structure to 1.95 Å. RESC5 harbors a dimethylarginine dimethylaminohydrolase-like (DDAH) fold. DDAH enzymes hydrolyze methylated arginine residues produced during protein degradation. However, RESC5 is missing two key catalytic DDAH residues and does bind DDAH substrate or product. Implications of the fold for RESC5 function are discussed. This structure provides the first structural view of an RESC protein.

59 BASIC BIOLOGICAL SCIENCES↗

Insights from a workplace SARS-CoV-2 specimen collection program, with genomes placed into global sequence phylogeny

In 2020, the Department of Energy established the National Virtual Biotechnology Laboratory (NVBL) to address key challenges associated with COVID-19. As part of that effort, Pacific Northwest National Laboratory (PNNL) established a capability to collect and analyze specimens from employees who self-reported symptoms consistent with the disease. During the spring and fall of 2021, 688 specimens were screened for SARS-CoV-2, with 64 (9.3%) testing positive using reverse-transcriptase quantitative PCR (RT-qPCR). Of these, 36 samples were released for research. All 36 positive samples released for research were sequenced and genotyped. Here, the relationship between patient age and viral load as measured by Ct values was measured and determined to be only weakly significant. Consensus sequences for each sample were placed into a global phylogeny and transmission dynamics were investigated, revealing that the closest relative for many samples was from outside of Washington state, indicating mixing of viral pools within geographic regions.

59 BASIC BIOLOGICAL SCIENCES↗

Crystal structure of the Arabidopsis SPIRAL2 C-terminal domain reveals a p80-Katanin-like domain

Epidermal cells of dark-grown plant seedlings reorient their cortical microtubule arrays in response to blue light from a net lateral orientation to a net longitudinal orientation with respect to the long axis of cells. The molecular mechanism underlying this microtubule array reorientation involves katanin, a microtubule severing enzyme, and a plant-specific microtubule associated protein called SPIRAL2. Katanin preferentially severs longitudinal microtubules, generating seeds that amplify the longitudinal array. Upon severing, SPIRAL2 binds nascent microtubule minus ends and limits their dynamics, thereby stabilizing the longitudinal array while the lateral array undergoes net depolymerization. To date, no experimental structural information is available for SPIRAL2 to help inform its mechanism. To gain insight into SPIRAL2 structure and function, we determined a 1.8 Å resolution crystal structure of the Arabidopsis thaliana SPIRAL2 C-terminal domain. The domain is composed of seven core α-helices, arranged in an α-solenoid. Amino-acid sequence conservation maps primarily to one face of the domain involving helices α1, α3, α5, and an extended loop, the α6-α7 loop. The domain fold is similar to, yet structurally distinct from the C-terminal domain of Ge-1 (an mRNA decapping complex factor involved in P-body localization) and, surprisingly, the C-terminal domain of the katanin p80 regulatory subunit. The katanin p80 C-terminal domain heterodimerizes with the MIT domain of the katanin p60 catalytic subunit, and in metazoans, binds the microtubule minus-end factors CAMSAP3 and ASPM. Structural analysis predicts that SPIRAL2 does not engage katanin p60 in a mode homologous to katanin p80. The SPIRAL2 structure highlights an interesting evolutionary convergence of domain architecture and microtubule minus-end localization between SPIRAL2 and katanin complexes, and establishes a foundation upon which structure-function analysis can be conducted to elucidate the role of this domain in the regulation of plant microtubule arrays.

59 BASIC BIOLOGICAL SCIENCES↗

Chaperone-tip adhesin complex is vital for synergistic activation of CFA/I fimbriae biogenesis

Colonization factor CFA/I defines the major adhesive fimbriae of enterotoxigenic Escherichia coli and mediates bacterial attachment to host intestinal epithelial cells. The CFA/I fimbria consists of a tip-localized minor adhesive subunit, CfaE, and thousands of copies of the major subunit CfaB polymerized into an ordered helical rod. Biosynthesis of CFA/I fimbriae requires the assistance of the periplasmic chaperone CfaA and outer membrane usher CfaC. Although the CfaE subunit is proposed to initiate the assembly of CFA/I fimbriae, how it performs this function remains elusive. Here, we report the establishment of an in vitro assay for CFA/I fimbria assembly and show that stabilized CfaA-CfaB and CfaA-CfaE binary complexes together with CfaC are sufficient to drive fimbria formation. The presence of both CfaA-CfaE and CfaC accelerates fimbria formation, while the absence of either component leads to linearized CfaB polymers in vitro. We further report the crystal structure of the stabilized CfaA-CfaE complex, revealing features unique for biogenesis of Class 5 fimbriae.

59 BASIC BIOLOGICAL SCIENCES↗

Neutralization profiles of HIV-1 viruses from the VRC01 Antibody Mediated Prevention (AMP) trials

The VRC01 Antibody Mediated Prevention (AMP) efficacy trials conducted between 2016 and 2020 showed for the first time that passively administered broadly neutralizing antibodies (bnAbs) could prevent HIV-1 acquisition against bnAb-sensitive viruses. HIV-1 viruses isolated from AMP participants who acquired infection during the study in the sub-Saharan African (HVTN 703/HPTN 081) and the Americas/European (HVTN 704/HPTN 085) trials represent a panel of currently circulating strains of HIV-1 and offer a unique opportunity to investigate the sensitivity of the virus to broadly neutralizing antibodies (bnAbs) being considered for clinical development. Pseudoviruses were constructed using envelope sequences from 218 individuals. The majority of viruses identified were clade B and C; with clades A, D, F and G and recombinants AC and BF detected at lower frequencies. We tested eight bnAbs in clinical development (VRC01, VRC07-523LS, 3BNC117, CAP256.25, PGDM1400, PGT121, 10–1074 and 10E8v4) for neutralization against all AMP placebo viruses (n = 76). Compared to older clade C viruses (1998–2010), the HVTN703/HPTN081 clade C viruses showed increased resistance to VRC07-523LS and CAP256.25. At a concentration of 1μg/ml (IC80), predictive modeling identified the triple combination of V3/V2-glycan/CD4bs-targeting bnAbs (10-1074/PGDM1400/VRC07-523LS) as the best against clade C viruses and a combination of MPER/V3/CD4bs-targeting bnAbs (10E8v4/10-1074/VRC07-523LS) as the best against clade B viruses, due to low coverage of V2-glycan directed bnAbs against clade B viruses. Overall, the AMP placebo viruses represent a valuable resource for defining the sensitivity of contemporaneous circulating viral strains to bnAbs and highlight the need to update reference panels regularly. Our data also suggests that combining bnAbs in passive immunization trials would improve coverage of global viruses.

60 APPLIED LIFE SCIENCES↗

OrthoPhylo

This software builds on PHAME developed at LANL to generate phylogenetic trees of bacterial whole genome sequences. Where PHAME uses whole genome alignments to generate informative sites to base tress on, PHAME-OuS annotates bacterial genes, identifies orthologous sequences, aligns related proteins, uses those alignments to inform transcript alignments, then builds trees with several methods. The first is a conventional gene concatenation and ML tree estemation method. The second attempts to reconcile gene tree with a unified species tree using quartets (ASTRAL). Both methods allow filtering of gene lists on number of species represented, length, and gappiness in order to tune noise-to-signal for tree estimation

Middlebrook, Earl↗

Genome-Wide Transcription Factor DNA Binding Sites and Gene Regulatory Networks in Clostridium thermocellum

Clostridium thermocellum is a thermophilic bacterium recognized for its natural ability to effectively deconstruct cellulosic biomass. While there is a large body of studies on the genetic engineering of this bacterium and its physiology to-date, there is limited knowledge in the transcriptional regulation in this organism and thermophilic bacteria in general. The study herein is the first report of a large-scale application of DNA-affinity purification sequencing (DAP-seq) to transcription factors (TFs) from a bacterium. We applied DAP-seq to > 90 TFs in C. thermocellum and detected genome-wide binding sites for 11 of them. We then compiled and aligned DNA binding sequences from these TFs to deduce the primary DNA-binding sequence motifs for each TF. These binding motifs are further validated with electrophoretic mobility shift assay (EMSA) and are used to identify individual TFs’ regulatory targets in C. thermocellum . Our results led to the discovery of novel, uncharacterized TFs as well as homologues of previously studied TFs including RexA-, LexA-, and LacI-type TFs. We then used these data to reconstruct gene regulatory networks for the 11 TFs individually, which resulted in a global network encompassing the TFs with some interconnections. As gene regulation governs and constrains how bacteria behave, our findings shed light on the roles of TFs delineated by their regulons, and potentially provides a means to enable rational, advanced genetic engineering of C. thermocellum and other organisms alike toward a desired phenotype.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-wide Transcription Factor DNA Binding Sites and Gene Regulatory Networks in Clostridium thermocellum

Clostridium thermocellum is a thermophilic bacterium recognized for its natural ability to effectively deconstruct cellulosic biomass. While there is a large body of studies on the genetic engineering of this bacterium and its physiology to-date, there is limited knowledge in the transcriptional regulation in this organism and thermophilic bacteria in general. The study herein is the first report of a high-throughput application of DNA-affinity purification sequencing (DAP-seq) to transcription factors (TFs) from a thermophile. We applied DAP-seq to >90 TFs in C. thermocellum and detected genome-wide binding sites for 11 of them. We then compiled and aligned DNA binding sequences from these TFs to deduce the primary DNA-binding sequence motifs for each TF. These binding motifs are further validated with electrophoretic mobility shift assay (EMSA) and are used to identify individual TFs’ regulatory targets in C. thermocellum. Our results led to the discovery of novel, uncharacterized TFs as well as homologues of previously studied TFs including RexA-, LexA- and LacI-type TFs. We then used these data to reconstruct gene regulatory networks for the 11 TFs individually, which resulted in a global network encompassing the TFs with some interconnections. As gene regulation governs and constrains how bacteria behave, our findings shed light on the roles of TFs delineated by their regulons, and potentially provides a means to enable rational, advanced genetic engineering of C. thermocellum and other organisms alike towards a desired phenotype.

09 BIOMASS FUELS↗

Sample Type Adaptable RNA depletion performance data [Slides]

Sequencing data was aligned to rRNA data sets QIIME_16S_MiDAS_4.8.1 and SILVA_138.1_LSURef_NR99 using BWA. RNA used was a composite (pool) of wastewater RNA samples collected at LANL between April 2022 and December 2022. RNA sample was split into four aliquots (STAR depletion and STAR depletion no-probe-control as well as RiboZero and RiboZero no-probe-control). All sequencing libraries were prepared from the depleted RNA samples using the same library prep method and sequenced on Illumina platforms.

59 BASIC BIOLOGICAL SCIENCES↗

CoverM: read alignment statistics for metagenomics

SUMMARY: Genome-centric analysis of metagenomic samples is a powerful method for understanding the function of microbial communities. Calculating read coverage is a central part of analysis, enabling differential coverage binning for recovery of genomes and estimation of microbial community composition. Coverage is determined by processing read alignments to reference sequences of either contigs or genomes. Per-reference coverage is typically calculated in an ad-hoc manner, with each software package providing its own implementation and specific definition of coverage. Here we present a unified software package CoverM which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner. It uses "Mosdepth arrays" for computational efficiency and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results. AVAILABILITY AND IMPLEMENTATION: CoverM is free software available at https://github.com/wwood/coverm. CoverM is implemented in Rust, with Python (https://github.com/apcamargo/pycoverm) and Julia (https://github.com/JuliaBinaryWrappers/CoverM_jll.jl) interfaces.

Aroney, Samuel T N↗

Identification of 2-Hydroxyacyl-CoA Synthases with High Acyloin Condensation Activity for Orthogonal One-Carbon Bioconversion

One-carbon (C1) compounds are emerging as cost-effective and potentially carbon-negative feedstocks for biomanufacturing, which require efficient, versatile metabolic platforms for the synthesis of value-added products. Synthetic formyl-CoA elongation (FORCE) pathways allow diverse product synthesis from C1 compounds via iterative C1 elongation, operating independently from the host metabolism with reduced engineering complexity and improved theoretical yields. However, a major bottleneck was identified as the suboptimal kinetics of the core C1–C1 condensation enzyme, 2-hydroxyacyl-CoA synthase (HACS), catalyzing the acyloin condensation reaction between formaldehyde and formyl-CoA. Furthermore, we used a combinatorial approach of bioprospecting and rational protein engineering to identify multiple HACS variants with significantly improved activities toward C1 substrates. Sequence and structure alignment of the active variants elucidated the key regions for the catalytic function, which were targeted for mutagenesis, leading to improved catalytic efficiency. In parallel, a consecutive round of bioprospecting for homologs with high similarity with active variants revealed a highly active HACS variant exhibiting up to 7-fold improvement in catalytic efficiency (k cat /K M ) and 14-fold improvement in the FORCE pathway flux in vivo compared to the previous reports. Upon further optimization of the downstream pathway, the orthogonal C1-to-product bioconversion system showed a metabolic flux of up to 700 μM glycolate OD –1 h –1 (2.1 mmol gDCW –1 h –1 ) and an industrially relevant glycolate titer, rate, and yield of 5.2 g L –1 (67.8 mM), 0.22 g L –1 h –1 , and 94% carbon yield, respectively.

2-hydroxyacyl-CoA synthase↗

PASTIS v0.1

PASTIS is a distributed-memory library for aligning large protein sequence sets and constructing protein similarity networks.

Buluc, Aydin↗