Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The prokaryote-to-eukaryote transition reflected in the evolution of the V/F/A-ATPase catalytic and proteolipid subunits

Changes in the primary and quarternary structure of vacuolar and archaeal type ATPases that accompany the prokaryote-to-eukaryote transition are analyzed. The gene encoding the vacuolar-type proteolipid of the V-ATPase from Giardia lamblia is reported. Giardia has a typical vacuolar ATPase as observed from the common motifs shared between its proteolipid subunit and other eukaryotic vacuolar ATPases, suggesting that the former enzyme works as a hydrolase in this primitive eukaryote. The phylogenetic analyses of the V-ATPase catalytic subunit and the front and back halves of the proteolipid subunit placed Giardia as the deepest branch within the eukaryotes. Our phylogenetic analysis indicated that at least two independent duplication and fusion events gave rise to the larger proteolipid type found in eukaryotes and in Methanococcus. The spatial distribution of the conserved residues among the vacuolar-type proteolipids suggest a zipper-type interaction among the transmembrane helices and surrounding subunits of the V-ATPase complex. Important residues involved in the function of the F-ATP synthase proteolipid have been replaced during evolution in the V-proteolipid, but in some cases retained in the archaeal A-ATPase. Their possible implication in the evolution of V/F/A-ATPases is discussed.

NASA Discipline Exobiology↗

The expression and function of the achaete-scute genes in Tribolium castaneum reveals conservation and variation in neural pattern formation and cell fate specification

The study of achaete-scute (ac/sc) genes has recently become a paradigm to understand the evolution and development of the arthropod nervous system. We describe the identification and characterization of the ac/sc genes in the coleopteran insect species Tribolium castaneum. We have identified two Tribolium ac/sc genes - achaete-scute homolog (Tc-ASH) a proneural gene and asense (Tc-ase) a neural precursor gene that reside in a gene complex. Focusing on the embryonic central nervous system we find that Tc-ASH is expressed in all neural precursors and the proneural clusters from which they segregate. Through RNAi and misexpression studies we show that Tc-ASH is necessary for neural precursor formation in Tribolium and sufficient for neural precursor formation in Drosophila. Comparison of the function of the Drosophila and Tribolium proneural ac/sc genes suggests that in the Drosophila lineage these genes have maintained their ancestral function in neural precursor formation and have acquired a new role in the fate specification of individual neural precursors. Furthermore, we find that Tc-ase is expressed in all neural precursors suggesting an important and conserved role for asense genes in insect nervous system development. Our analysis of the Tribolium ac/sc genes indicates significant plasticity in gene number, expression and function, and implicates these modifications in the evolution of arthropod neural development.

Non-NASA Center↗

Expression of the ctenophore Brain Factor 1 forkhead gene ortholog (ctenoBF-1) mRNA is restricted to the presumptive mouth and feeding apparatus: implications for axial organization in the Metazoa

Ctenophores are thoroughly modern animals whose ancestors are derived from a separate evolutionary branch than that of other eumetazoans. Their major longitudinal body axis is the oral-aboral axis. An apical sense organ, called the apical organ, is located at the aboral pole and contains a highly innervated statocyst and photodetecting cells. The apical organ integrates sensory information and controls the locomotory apparatus of ctenophores, the eight longitudinal rows of ctene/comb plates. In an effort to understand the developmental and evolutionary organization of axial properties of ctenophores we have isolated a forkhead gene from the Brain Factor 1 (BF-1) family. This gene, ctenoBF-1, is the first full-length nuclear gene reported from ctenophores. This makes ctenophores the most basal metazoan (to date) known to express definitive forkhead class transcription factors. Orthologs of BF-1 in vertebrates, Drosophila, and Caenorhabditis elegans are expressed in anterior neural structures. Surprisingly, in situ hybridizations with ctenoBF-1 antisense riboprobes show that this gene is not expressed in the apical organ of ctenophores. CtenoBF-1 is expressed prior to first cleavage. Transcripts become localized to the aboral pole by the 8-cell stage and are inherited by ectodermal micromeres generated from this region at the 16- and 32-cell stages. Expression in subsets of these cells persists and is seen around the edge of the blastopore (presumptive mouth) and in distinct ectodermal regions along the tentacular poles. Following gastrulation, stomodeal expression begins to fade and intense staining becomes restricted to two distinct domains in each tentacular feeding apparatus. We suggest that the apical organ is not homologous to the brain of bilaterians but that the oral pole of ctenophores corresponds to the anterior pole of bilaterian animals.

Non-NASA Center↗

Classification of bacterial plasmid and chromosome derived sequences using machine learning

Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.

59 BASIC BIOLOGICAL SCIENCES↗

SCOPe: improvements to the structural classification of proteins – extended database to facilitate variant interpretation and machine learning

Abstract The Structural Classification of Proteins—extended (SCOPe, https://scop.berkeley.edu) knowledgebase aims to provide an accurate, detailed, and comprehensive description of the structural and evolutionary relationships amongst the majority of proteins of known structure, along with resources for analyzing the protein structures and their sequences. Structures from the PDB are divided into domains and classified using a combination of manual curation and highly precise automated methods. In the current release of SCOPe, 2.08, we have developed search and display tools for analysis of genetic variants we mapped to structures classified in SCOPe. In order to improve the utility of SCOPe to automated methods such as deep learning classifiers that rely on multiple alignment of sequences of homologous proteins, we have introduced new machine-parseable annotations that indicate aberrant structures as well as domains that are distinguished by a smaller repeat unit. We also classified structures from 74 of the largest Pfam families not previously classified in SCOPe, and we improved our algorithm to remove N- and C-terminal cloning, expression and purification sequences from SCOPe domains. SCOPe 2.08-stable classifies 106 976 PDB entries (about 60% of PDB entries).

59 BASIC BIOLOGICAL SCIENCES↗

Functional characteristics of the calcium modulated proteins seen from an evolutionary perspective

We have constructed dendrograms relating 173 EF-hand proteins of known amino acid sequence. We aligned all of these proteins by their EF-hand domains, omitting interdomain regions. Initial dendrograms were computed by minimum mutation distance methods. Using these as starting points, we determined the best dendrogram by the method of maximum parsimony, scored by minimum mutation distance. We identified 14 distinct subfamilies as well as 6 unique proteins that are perhaps the sole representatives of other subfamilies. This information is given in tabular form. Within subfamilies one can easily align interdomain regions. The resulting dendrograms are very similar to those computed using domains only. Dendrograms constructed using pairs of domains show general congruence. However, there are enough exceptions to caution against an overly simple scheme in which one pair of gene duplications leads from one domain precurser to a four domain prototype from which all other forms evolved. The ability to bind calcium was lost and acquired several times during evolution. The distribution of introns does not conform to the dendrogram based on amino acid sequences. The rates of evolution appear to be much slower within subfamilies, especially within calmodulin, than those prior to the definition of subfamily.

Kretsinger, R. H.↗

Analysis of genomic signatures associated with Variovorax endosphere colonization

This repository contains the analysis code and supporting datasets associated with the study “Genomic signatures in Variovorax enabling colonization of the Populus endosphere.” Beals DG, Carper DL, Hochanadel LH, Jawdy SS, Klingeman DM, Piatkowski BT, Weston DJ, Doktycz MJ, Pelletier DA. 2026. Genomic signatures in Variovorax enabling colonization of the Populus endosphere. mSystems 11:e01605-25. https://doi.org/10.1128/msystems.01605-25 The scripts are organized sequentially (01–07) and document the workflows used for: Sequence-read alignment and feature counting Orthogroup and KEGG Ortholog annotation Count normalization Statistical analysis and aggregation Generation of manuscript figures and tables Repository contents The uncompressed files are the finalized, formatted datasets used to generate the figures and tables reported in the study, including the supplemental CSV files referenced in the manuscript. The accompanying ZIP archive contains the complete codebase and example data_input/ and data_output/ directories illustrating the organization and execution of the analytical workflow. Individual scripts identify the corresponding manuscript analyses and figure panels. Raw sequencing data Raw sequencing reads are available through the NCBI Sequence Read Archive under BioProject accession PRJNA1322484.

Beals, Delaney [ORNL] (ORCID:0000000306274574)↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗

Tandem repeats in giant archaeal Borg elements undergo rapid evolution and create new intrinsically disordered regions in proteins

Borgs are huge, linear extrachromosomal elements associated with anaerobic methane-oxidizing archaea. Striking features of Borg genomes are pervasive tandem direct repeat (TR) regions. Here, we present six new Borg genomes and investigate the characteristics of TRs in all ten complete Borg genomes. We find that TR regions are rapidly evolving, recently formed, arise independently, and are virtually absent in host Methanoperedens genomes. Flanking partial repeats and A-enriched character constrain the TR formation mechanism. TRs can be in intergenic regions, where they might serve as regulatory RNAs, or in open reading frames (ORFs). TRs in ORFs are under very strong selective pressure, leading to perfect amino acid TRs (aaTRs) that are commonly intrinsically disordered regions. Proteins with aaTRs are often extracellular or membrane proteins, and functionally similar or homologous proteins often have aaTRs composed of the same amino acids. We propose that Borg aaTR-proteins functionally diversify Methanoperedens and all TRs are crucial for specific Borg–host associations and possibly cospeciation.

59 BASIC BIOLOGICAL SCIENCES↗

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Tracking ebolavirus genomic drift with a resequencing microarray

Filoviruses are emerging pathogens that cause acute fever with high fatality rate and present a global public health threat. During the 2013–2016 Ebola virus outbreak, genome sequencing allowed the study of virus evolution, mutations affecting pathogenicity and infectivity, and tracing the viral spread. In 2018, early sequence identification of the Ebolavirus as EBOV in the Democratic Republic of the Congo supported the use of an Ebola virus vaccine. However, field-deployable sequencing methods are needed to enable a rapid public health response. Resequencing microarrays (RMA) are a targeted method to obtain genomic sequence on clinical specimens rapidly, and sensitively, overcoming the need for extensive bioinformatic analysis. This study presents the design and initial evaluation of an ebolavirus resequencing microarray (Ebolavirus-RMA) system for sequencing the major genomic regions of four Ebolaviruses that cause disease in humans. The design of the Ebolavirus-RMA system is described and evaluated by sequencing repository samples of three Ebolaviruses and two EBOV variants. The ability of the system to identify genetic drift in a replicating virus was achieved by sequencing the ebolavirus glycoprotein gene in a recombinant virus cultured under pressure from a neutralizing antibody. Comparison of the Ebolavirus-RMA results to the Genbank database sequence file with the accession number given for the source RNA and Ebolavirus-RMA results compared to Next Generation Sequence results of the same RNA samples showed up to 99% agreement.

59 BASIC BIOLOGICAL SCIENCES↗

Structure of the T. brucei kinetoplastid RNA editing substrate-binding complex core component, RESC5

Kinetoplastid protists such as Trypanosoma brucei undergo an unusual process of mitochondrial uridine (U) insertion and deletion editing termed kinetoplastid RNA editing (kRNA editing). This extensive form of editing, which is mediated by guide RNAs (gRNAs), can involve the insertion of hundreds of Us and deletion of tens of Us to form a functional mitochondrial mRNA transcript. kRNA editing is catalyzed by the 20 S editosome/RECC. However, gRNA directed, processive editing requires the RNA editing substrate binding complex (RESC), which is comprised of 6 core proteins, RESC1-RESC6. To date there are no structures of RESC proteins or complexes and because RESC proteins show no homology to proteins of known structure, their molecular architecture remains unknown. RESC5 is a key core component in forming the foundation of the RESC complex. To gain insight into the RESC5 protein we performed biochemical and structural studies. We show that RESC5 is monomeric and we report the T . brucei RESC5 crystal structure to 1.95 Å. RESC5 harbors a dimethylarginine dimethylaminohydrolase-like (DDAH) fold. DDAH enzymes hydrolyze methylated arginine residues produced during protein degradation. However, RESC5 is missing two key catalytic DDAH residues and does bind DDAH substrate or product. Implications of the fold for RESC5 function are discussed. This structure provides the first structural view of an RESC protein.

59 BASIC BIOLOGICAL SCIENCES↗

Insights from a workplace SARS-CoV-2 specimen collection program, with genomes placed into global sequence phylogeny

In 2020, the Department of Energy established the National Virtual Biotechnology Laboratory (NVBL) to address key challenges associated with COVID-19. As part of that effort, Pacific Northwest National Laboratory (PNNL) established a capability to collect and analyze specimens from employees who self-reported symptoms consistent with the disease. During the spring and fall of 2021, 688 specimens were screened for SARS-CoV-2, with 64 (9.3%) testing positive using reverse-transcriptase quantitative PCR (RT-qPCR). Of these, 36 samples were released for research. All 36 positive samples released for research were sequenced and genotyped. Here, the relationship between patient age and viral load as measured by Ct values was measured and determined to be only weakly significant. Consensus sequences for each sample were placed into a global phylogeny and transmission dynamics were investigated, revealing that the closest relative for many samples was from outside of Washington state, indicating mixing of viral pools within geographic regions.

59 BASIC BIOLOGICAL SCIENCES↗

Crystal structure of the Arabidopsis SPIRAL2 C-terminal domain reveals a p80-Katanin-like domain

Epidermal cells of dark-grown plant seedlings reorient their cortical microtubule arrays in response to blue light from a net lateral orientation to a net longitudinal orientation with respect to the long axis of cells. The molecular mechanism underlying this microtubule array reorientation involves katanin, a microtubule severing enzyme, and a plant-specific microtubule associated protein called SPIRAL2. Katanin preferentially severs longitudinal microtubules, generating seeds that amplify the longitudinal array. Upon severing, SPIRAL2 binds nascent microtubule minus ends and limits their dynamics, thereby stabilizing the longitudinal array while the lateral array undergoes net depolymerization. To date, no experimental structural information is available for SPIRAL2 to help inform its mechanism. To gain insight into SPIRAL2 structure and function, we determined a 1.8 Å resolution crystal structure of the Arabidopsis thaliana SPIRAL2 C-terminal domain. The domain is composed of seven core α-helices, arranged in an α-solenoid. Amino-acid sequence conservation maps primarily to one face of the domain involving helices α1, α3, α5, and an extended loop, the α6-α7 loop. The domain fold is similar to, yet structurally distinct from the C-terminal domain of Ge-1 (an mRNA decapping complex factor involved in P-body localization) and, surprisingly, the C-terminal domain of the katanin p80 regulatory subunit. The katanin p80 C-terminal domain heterodimerizes with the MIT domain of the katanin p60 catalytic subunit, and in metazoans, binds the microtubule minus-end factors CAMSAP3 and ASPM. Structural analysis predicts that SPIRAL2 does not engage katanin p60 in a mode homologous to katanin p80. The SPIRAL2 structure highlights an interesting evolutionary convergence of domain architecture and microtubule minus-end localization between SPIRAL2 and katanin complexes, and establishes a foundation upon which structure-function analysis can be conducted to elucidate the role of this domain in the regulation of plant microtubule arrays.

59 BASIC BIOLOGICAL SCIENCES↗

Chaperone-tip adhesin complex is vital for synergistic activation of CFA/I fimbriae biogenesis

Colonization factor CFA/I defines the major adhesive fimbriae of enterotoxigenic Escherichia coli and mediates bacterial attachment to host intestinal epithelial cells. The CFA/I fimbria consists of a tip-localized minor adhesive subunit, CfaE, and thousands of copies of the major subunit CfaB polymerized into an ordered helical rod. Biosynthesis of CFA/I fimbriae requires the assistance of the periplasmic chaperone CfaA and outer membrane usher CfaC. Although the CfaE subunit is proposed to initiate the assembly of CFA/I fimbriae, how it performs this function remains elusive. Here, we report the establishment of an in vitro assay for CFA/I fimbria assembly and show that stabilized CfaA-CfaB and CfaA-CfaE binary complexes together with CfaC are sufficient to drive fimbria formation. The presence of both CfaA-CfaE and CfaC accelerates fimbria formation, while the absence of either component leads to linearized CfaB polymers in vitro. We further report the crystal structure of the stabilized CfaA-CfaE complex, revealing features unique for biogenesis of Class 5 fimbriae.

59 BASIC BIOLOGICAL SCIENCES↗

Neutralization profiles of HIV-1 viruses from the VRC01 Antibody Mediated Prevention (AMP) trials

The VRC01 Antibody Mediated Prevention (AMP) efficacy trials conducted between 2016 and 2020 showed for the first time that passively administered broadly neutralizing antibodies (bnAbs) could prevent HIV-1 acquisition against bnAb-sensitive viruses. HIV-1 viruses isolated from AMP participants who acquired infection during the study in the sub-Saharan African (HVTN 703/HPTN 081) and the Americas/European (HVTN 704/HPTN 085) trials represent a panel of currently circulating strains of HIV-1 and offer a unique opportunity to investigate the sensitivity of the virus to broadly neutralizing antibodies (bnAbs) being considered for clinical development. Pseudoviruses were constructed using envelope sequences from 218 individuals. The majority of viruses identified were clade B and C; with clades A, D, F and G and recombinants AC and BF detected at lower frequencies. We tested eight bnAbs in clinical development (VRC01, VRC07-523LS, 3BNC117, CAP256.25, PGDM1400, PGT121, 10–1074 and 10E8v4) for neutralization against all AMP placebo viruses (n = 76). Compared to older clade C viruses (1998–2010), the HVTN703/HPTN081 clade C viruses showed increased resistance to VRC07-523LS and CAP256.25. At a concentration of 1μg/ml (IC80), predictive modeling identified the triple combination of V3/V2-glycan/CD4bs-targeting bnAbs (10-1074/PGDM1400/VRC07-523LS) as the best against clade C viruses and a combination of MPER/V3/CD4bs-targeting bnAbs (10E8v4/10-1074/VRC07-523LS) as the best against clade B viruses, due to low coverage of V2-glycan directed bnAbs against clade B viruses. Overall, the AMP placebo viruses represent a valuable resource for defining the sensitivity of contemporaneous circulating viral strains to bnAbs and highlight the need to update reference panels regularly. Our data also suggests that combining bnAbs in passive immunization trials would improve coverage of global viruses.

60 APPLIED LIFE SCIENCES↗

The McIDAS system

The man-computer interactive data access system (McIDAS) hardware and software design is outlined, with emphasis on meteorological applications. The McIDAS system features a flexible digital image enhancement device. Theory of operation, the McIDAS language, the McIDAS executive monitor, image acquisition, image parameters, image display, graphics, image navigation, and image massaging are explained. The WINDCO/CLDHGT system, developed to align time sequences of geosynchronous satellite pictures and to determine the motion and height of selected cloud targets, is also described.

Smith, E. A.↗

The potato virus X TGBp2 protein association with the endoplasmic reticulum plays a role in but is not sufficient for viral cell-to-cell movement

Potato virus X (PVX) TGBp1, TGBp2, TGBp3, and coat protein are required for virus cell-to-cell movement. Plasmids expressing GFP fused to TGBp2 were bombarded to leaf epidermal cells and GFP:TGBp2 moved cell to cell in Nicotiana benthamiana leaves but not in Nicotiana tabacum leaves. GFP:TGBp2 movement was observed in TGBp1-transgenic N. tabacum, indicating that TGBp2 requires TGBp1 to promote its movement in N. tabacum. In this study, GFP:TGBp2 was detected in a polygonal pattern that resembles the endoplasmic reticulum (ER) network. Amino acid sequence analysis revealed TGBp2 has two putative transmembrane domains. Two mutations separately introduced into the coding sequences encompassing the putative transmembrane domains within the GFP:TGBp2 plasmids and PVX genome, disrupted membrane binding of GFP:TGBp2, inhibited GFP:TGBp2 movement in N. benthamiana and TGBp1-expressing N. tabacum, and inhibited PVX movement. A third mutation, lying outside the transmembrane domains, had no effect on GFP:TGBp2 ER association or movement in N. benthamiana but inhibited GFP:TGBp2 movement in TGBp1-expressing N. tabacum and PVX movement in either Nicotiana species. Thus, ER association of TGBp2 may be required but not be sufficient for virus movement. TGBp2 likely provides an activity for PVX movement beyond ER association.

NASA Discipline Plant Biology↗