Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data lineage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Delineating the mechanism of anti-Lassa virus GPC-A neutralizing antibodies

Lassa virus (LASV) is the etiologic agent of Lassa Fever, a hemorrhagic disease that is endemic to West Africa. During LASV infection, LASV glycoprotein (GP) engages with multiple host receptors for cell entry. Neutralizing antibodies against GP are rare and principally target quaternary epitopes displayed only on the metastable, pre-fusion conformation of GP. Currently, the structural features of the neutralizing GPC-A antibody competition group are understudied. Structures of two GPC-A antibodies presented here demonstrate that they bind the side of the pre-fusion GP trimer, bridging the GP1 and GP2 subunits. Complementary biochemical analyses indicate that antibody 25.10C, which is broadly specific, neutralizes by inhibiting binding of the endosomal receptor LAMP1 and also by blocking membrane fusion. The other GPC-A antibody, 36.1F, which is lineage-specific, prevents LAMP1 association only. These data illuminate a site of vulnerability on LASV GP and will guide efforts to elicit broadly reactive therapeutics and vaccines.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating biological nitrogen fixation via single-cell transcriptomics

The extensive use of nitrogen fertilizers has detrimental environmental consequences, and it is essential for society to explore sustainable alternatives. One promising avenue is engineering root nodule symbiosis, a naturally occurring process in certain plant species within the nitrogen-fixing clade, into non-leguminous crops. Advancements in single-cell transcriptomics provide unprecedented opportunities to dissect the molecular mechanisms underlying root nodule symbiosis at the cellular level. This review summarizes key findings from single-cell studies in Medicago truncatula, Lotus japonicus, and Glycine max. We highlight how these studies address fundamental questions about the development of root nodule symbiosis, including the following findings: (i) single-cell transcriptomics has revealed a conserved transcriptional program in root hair and cortical cells during rhizobial infection, suggesting a common infection pathway across legume species; (ii) characterization of determinate and indeterminate nodules using single-cell technologies supports the compartmentalization of nitrogen fixation, assimilation, and transport into distinct cell populations; (iii) single-cell transcriptomics data have enabled the identification of novel root nodule symbiosis genes and provided new approaches for prioritizing candidate genes for functional characterization; and (iv) trajectory inference and RNA velocity analyses of single-cell transcriptomics data have allowed the reconstruction of cellular lineages and dynamic transcriptional states during root nodule symbiosis.

Lotus japonicus↗

Whole genome resequencing data from a collection of Clostridium Thermocellum strains

Clostridium thermocellum is an anaerobic thermophilic bacterium that natively ferments cellulose to ethanol and organic acids. This data set is a collection of whole genome resequencing data for several hundred strains of Clostridium thermocellum. It includes strains that have been engineered to increase ethanol production, strains that have been engineered to understand microbial physiology, and strains that have been adapted for desired phenotypes including increased ethanol tolerance. Resequencing data consists of paired Illumina reads, 100-150 bp on each end, with a ~500 bp insert size. One data file containing raw Illumina data (interleaved) is available for each strain. We also provide data describing the mutations identified in each strain, and distinguish between inherited and newly observed mutations. In addition to resequencing data, we also provide metadata describing the lineage of each strain, and any targeted genetic modifications.

resequencing bio energy fermentation↗

SARS-CoV-2 wastewater variant surveillance: pandemic response leveraging FDA’s GenomeTrakr network

ABSTRACT Wastewater surveillance has emerged as a crucial public health tool for population-level pathogen surveillance. Supported by funding from the American Rescue Plan Act of 2021, the FDA‘s genomic epidemiology program, GenomeTrakr, was leveraged to sequence SARS-CoV-2 from wastewater sites across the United States. This initiative required the evaluation, optimization, development, and publication of new methods and analytical tools spanning sample collection through variant analyses. Version-controlled protocols for each step of the process were developed and published on protocols.io. A custom data analysis tool and a publicly accessible dashboard were built to facilitate real-time visualization of the collected data, focusing on the relative abundance of SARS-CoV-2 variants and sub-lineages across different samples and sites throughout the project. From September 2021 through June 2023, a total of 3,389 wastewater samples were collected, with 2,517 undergoing sequencing and submission to NCBI under the umbrella BioProject,PRJNA757291. Sequence data were released with explicit quality control (QC) tags on all sequence records, communicating our confidence in the quality of data. Variant analysis revealed wide circulation of Delta in the fall of 2021 and captured the sweep of Omicron and subsequent diversification of this lineage through the end of the sampling period. This project successfully achieved two important goals for the FDA’s GenomeTrakr program: first, contributing timely genomic data for the SARS-CoV-2 pandemic response, and second, establishing both capacity and best practices for culture-independent, population-level environmental surveillance for other pathogens of interest to the FDA. IMPORTANCE This paper serves two primary objectives. First, it summarizes the genomic and contextual data collected during a Covid-19 pandemic response project, which utilized the FDA’s laboratory network, traditionally employed for sequencing foodborne pathogens, for sequencing SARS-CoV-2 from wastewater samples. Second, it outlines best practices for gathering and organizing population-level next generation sequencing (NGS) data collected for culture-free, surveillance of pathogens sourced from environmental samples.

Microbiology↗

yProv4ML: Effortless provenance tracking for machine learning systems

The rapid growth in interest in deep learning and foundation models (FMs) in particular, has attracted the attention of a diverse range of researchers thanks to their generalization ability. However, the advent of these techniques has also brought to light the lack of transparency and rigor in the way development is pursued. In particular, the inability to determine the number of epochs and other hyperparameters in advance presents challenges in identifying the best model. To address this challenge, machine learning frameworks such as MLFlow can automate the collection of this type of information. However, these tools capture data using proprietary formats and pose little attention to lineage. This paper proposes yProv4ML, a framework that captures provenance information generated during machine learning processes in PROV-JSON format, with minimal code modification.

Machine learning↗

Phytoplankton exudates and lysates support distinct microbial consortia with specialized metabolic and ecophysiological traits

Marine dissolved organic matter, which originates from phytoplankton, holds as much carbon as Earth’s atmosphere; yet, the biological processes governing its fate are primarily studied under idealized laboratory conditions or through indirect measures such as genome sequencing. In this work, we used isotope labeling to directly quantify uptake of complex carbon pools from the two primary sources of marine organic carbon (diatoms and cyanobacteria) by a natural microbial community. Furthermore, our data show that carbon pools are partitioned into distinct microbial lineages whose physiological properties and resource acquisition strategies match the chemical nature of their preferred substrates. Our results provide ecological and functional insights into the patterns of microbial community structure changes that occur during marine phytoplankton blooms.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms↗

Ancient Rapid Radiation Explains Most Conflicts Among Gene Trees and Well-Supported Phylogenomic Trees of Nostocalean Cyanobacteria

Prokaryotic genomes are often considered to be mosaics of genes that do not necessarily share the same evolutionary history due to widespread horizontal gene transfers (HGTs). Consequently, representing evolutionary relationships of prokaryotes as bifurcating trees has long been controversial. However, studies reporting conflicts among gene trees derived from phylogenomic data sets have shown that these conflicts can be the result of artifacts or evolutionary processes other than HGT, such as incomplete lineage sorting, low phylogenetic signal, and systematic errors due to substitution model misspecification. Here, we present the results of an extensive exploration of phylogenetic conflicts in the cyanobacterial order Nostocales, for which previous studies have inferred strongly supported conflicting relationships when using different concatenated phylogenomic data sets. We found that most of these conflicts are concentrated in deep clusters of short internodes of the Nostocales phylogeny, where the great majority of individual genes have low resolving power. We then inferred phylogenetic networks to detect HGT events while also accounting for incomplete lineage sorting. Our results indicate that most conflicts among gene trees are likely due to incomplete lineage sorting linked to an ancient rapid radiation, rather than to HGTs. Moreover, the short internodes of this radiation fit the expectations of the anomaly zone, i.e., a region of the tree parameter space where a species tree is discordant with its most likely gene tree. In this work, we demonstrated that concatenation of different sets of loci can recover up to 17 distinct and well-supported relationships within the putative anomaly zone of Nostocales, corresponding to the observed conflicts among well-supported trees based on concatenated data sets from previous studies. Our findings highlight the important role of rapid radiations as a potential cause of strongly conflicting phylogenetic relationships when using phylogenomic data sets of bacteria. We propose that polytomies may be the most appropriate phylogenetic representation of these rapid radiations that are part of anomaly zones, especially when all possible genomic markers have been considered to infer these phylogenies.

59 BASIC BIOLOGICAL SCIENCES↗

Complete genomes of Asgard archaea reveal diverse integrated and mobile genetic elements

Asgard archaea are of great interest as the progenitors of Eukaryotes, but little is known about the mobile genetic elements (MGEs) that may shape their ongoing evolution. Here, we describe MGEs that replicate in Atabeyarchaeia, a wetland Asgard archaea lineage represented by two complete genomes. We used soil depth–resolved population metagenomic data sets to track 18 MGEs for which genome structures were defined and precise chromosome integration sites could be identified for confident host linkage. Additionally, we identified a complete 20.67 kbp circular plasmid and two family-level groups of viruses linked to Atabeyarchaeia, via CRISPR spacer targeting. Closely related 40 kbp viruses possess a hypervariable genomic region encoding combinations of specific genes for small cysteine-rich proteins structurally similar to restriction-homing endonucleases. One 10.9 kbp integrative conjugative element (ICE) integrates genomically into theAtabeyarchaeum deiterrae-1chromosome and has a 2.5 kbp circularizable element integrated within it. The 10.9 kbp ICE encodes an expressed Type IIG restriction-modification system with a sequence specificity matching an active methylation motif identified by Pacific Biosciences (PacBio) high-accuracy long-read (HiFi) metagenomic sequencing. Restriction-modification of Atabeyarchaeia differs from that of another coexisting Asgard archaea, Freyarchaeia, which has few identified MGEs but possesses diverse defense mechanisms, including DISARM and Hachiman, not found in Atabeyarchaeia. Overall, defense systems and methylation mechanisms of Asgard archaea likely modulate their interactions with MGEs, and integration/excision and copy number variation of MGEs in turn enable host genetic versatility.

Biochemistry & Molecular Biology↗

Intrahost SARS-CoV-2 k-mer Identification Method (iSKIM) for Rapid Detection of Mutations of Concern Reveals Emergence of Global Mutation Patterns

Despite unprecedented global sequencing and surveillance of SARS-CoV-2, timely identification of the emergence and spread of novel variants of concern (VoCs) remains a challenge. Several million raw genome sequencing runs are now publicly available. We sought to survey these datasets for intrahost variation to study emerging mutations of concern. We developed iSKIM (“intrahost SARS-CoV-2 k-mer identification method”) to relatively quickly and efficiently screen the many SARS-CoV-2 datasets to identify intrahost mutations belonging to lineages of concern. Certain mutations surged in frequency as intrahost minor variants just prior to, or while lineages of concern arose. The Spike N501Y change common to several VoCs was found as a minor variant in 834 samples as early as October 2020. This coincides with the timing of the first detected samples with this mutation in the Alpha/B.1.1.7 and Beta/B.1.351 lineages. Using iSKIM, we also found that Spike L452R was detected as an intrahost minor variant as early as September 2020, prior to the observed rise of the Epsilon/B.1.429/B.1.427 lineages in late 2020. iSKIM rapidly screens for mutations of interest in raw data, prior to genome assembly, and can be used to detect increases in intrahost variants, potentially providing an early indication of novel variant spread.

60 APPLIED LIFE SCIENCES↗

Combining biomarker and virus phylogenetic models improves HIV-1 epidemiological source identification

To identify and stop active HIV transmission chains new epidemiological techniques are needed. Here, we describe the development of a multi-biomarker augmentation to phylogenetic inference of the underlying transmission history in a local population. HIV biomarkers are measurable biological quantities that have some relationship to the amount of time someone has been infected with HIV. To train our model, we used five biomarkers based on real data from serological assays, HIV sequence data, and target cell counts in longitudinally followed, untreated patients with known infection times. The biomarkers were modeled with a mixed effects framework to allow for patient specific variation and general trends, and fit to patient data using Markov Chain Monte Carlo (MCMC) methods. Subsequently, the density of the unobserved infection time conditional on observed biomarkers were obtained by integrating out the random effects from the model fit. This probabilistic information about infection times was incorporated into the likelihood function for the transmission history and phylogenetic tree reconstruction, informed by the HIV sequence data. To critically test our methodology, we developed a coalescent-based simulation framework that generates phylogenies and biomarkers given a specific or general transmission history. Testing on many epidemiological scenarios showed that biomarker augmented phylogenetics can reach 90% accuracy under idealized situations. Under realistic within-host HIV-1 evolution, involving substantial within-host diversification and frequent transmission of multiple lineages, the average accuracy was at about 50% in transmission clusters involving 5–50 hosts. Realistic biomarker data added on average 16 percentage points over using the phylogeny alone. Using more biomarkers improved the performance. Shorter temporal spacing between transmission events and increased transmission heterogeneity reduced reconstruction accuracy, but larger clusters were not harder to get right. More sequence data per infected host also improved accuracy. We show that the method is robust to incomplete sampling and that adding biomarkers improves reconstructions of real HIV-1 transmission histories. The technology presented here could allow for better prevention programs by providing data for locally informed and tailored strategies.

60 APPLIED LIFE SCIENCES↗

Mitochondrial Genomes of the United States Distribution of Gray Fox (Urocyon cinereoargenteus) Reveal a Major Phylogeographic Break at the Great Plains Suture Zone

We examined phylogeographic structure in gray fox (Urocyon cinereoargenteus) across the United States to identify the location of secondary contact zone(s) between eastern and western lineages and investigate the possibility of additional cryptic intraspecific divergences. We generated and analyzed complete mitochondrial genome sequence data from 75 samples and partial control region mitochondrial DNA sequences from 378 samples to investigate levels of genetic diversity and structure through population- and individual-based analyses including estimates of divergence (FST and SAMOVA), median joining networks, and phylogenies. We used complete mitochondrial genomes to infer phylogenetic relationships and date divergence times of major lineages of Urocyon in the United States. Despite broad-scale sampling, we did not recover additional major lineages of Urocyon within the United States, but identified a deep east-west split (~0.8 million years) with secondary contact at the Great Plains Suture Zone and confirmed the Channel Island fox (Urocyon littoralis) is nested within U. cinereoargenteus. Genetic diversity declined at northern latitudes in the eastern United States, a pattern concordant with post-glacial recolonization and range expansion. Beyond the east-west divergence, morphologically-based subspecies did not form monophyletic groups, though unique haplotypes were often geographically limited. Gray foxes in the United States displayed a deep, cryptic divergence suggesting taxonomic revision is needed. Secondary contact at a common phylogeographic break, the Great Plains Suture Zone, where environmental variables show a sharp cline, suggests ongoing evolutionary processes may reinforce this divergence. Follow-up study with nuclear markers should investigate whether hybridization is occurring along the suture zone and characterize contemporary population structure to help identify conservation units. Comparative work on other wide-ranging carnivores in the region should test whether similar evolutionary patterns and processes are occurring.

54 ENVIRONMENTAL SCIENCES↗

Pangenomics reveals alternative environmental lifestyles among chlamydiae

Chlamydiae are highly successful strictly intracellular bacteria associated with diverse eukaryotic hosts. Here we analyzed metagenome-assembled genomes of the “Genomes from Earth’s Microbiomes” initiative from diverse environmental samples, which almost double the known phylogenetic diversity of the phylum and facilitate a highly resolved view at the chlamydial pangenome. Chlamydiae are defined by a relatively large core genome indicative of an intracellular lifestyle, and a highly dynamic accessory genome of environmental lineages. We observe chlamydial lineages that encode enzymes of the reductive tricarboxylic acid cycle and for light-driven ATP synthesis. We show a widespread potential for anaerobic energy generation through pyruvate fermentation or the arginine deiminase pathway, and we add lineages capable of molecular hydrogen production. Genome-informed analysis of environmental distribution revealed lineage-specific niches and a high abundance of chlamydiae in some habitats. Together, our data provide an extended perspective of the variability of chlamydial biology and the ecology of this phylum of intracellular microbes.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic and Antigenic Characterization of an Expanding H3 Influenza A Virus Clade in U.S. Swine Visualized by Nextstrain

Defining factors that influence spatial and temporal patterns of influenza A virus (IAV) is essential to inform vaccine strain selection and strategies to reduce the spread of potentially zoonotic swine-origin IAV. The relative frequency of detection of the H3 phylogenetic clade 1990.4.a (colloquially known as C-IVA) in U.S. swine declined to 7% in 2017 but increased to 32% in 2019. We conducted phylogenetic and phenotypic analyses to determine putative mechanisms associated with increased detection. We created an implementation of Nextstrain to visualize the emergence, spatial spread, and genetic evolution of H3 IAV in swine, identifying two C-IVA clades that emerged in 2017 and cocirculated in multiple U.S. states. Phylodynamic analysis of the hemagglutinin (HA) gene documented low relative genetic diversity from 2017 to 2019, suggesting clonal expansion. The major H3 C-IVA clade contained an N156H amino acid substitution, but hemagglutination inhibition (HI) assays demonstrated no significant antigenic drift. The minor HA clade was paired with the neuraminidase (NA) clade N2-2002B prior to 2016 but acquired and maintained an N2-2002A in 2016, resulting in a loss of antigenic cross-reactivity between N2-2002B- and -2002A-containing H3N2 strains. The major C-IVA clade viruses acquired a nucleoprotein (NP) of the H1N1pdm09 lineage through reassortment in the replacement of the North American swine-lineage NP. Instead of genetic or antigenic diversity within the C-IVA HA, our data suggest that population immunity to H3 2010.1 along with the antigenic diversity of the NA and the acquisition of the H1N1pdm09 NP gene likely explain the reemergence and transmission of C-IVA H3N2 in swine.

59 BASIC BIOLOGICAL SCIENCES↗

Non-photosynthetic lineages sibling to Cyanobacteria associate with eukaryotes in the open ocean

Margulisbacteria are elusive uncultivated bacteria that have illuminated evolutionary transitions in the progenitor of Cyanobacteria, the latter being a critically important phylum that underpins oxygenic photosynthesis. The non-photosynthetic Margulisbacteria were discovered in a sulfidic spring and later in other habitats. Currently, this candidate phylum partitions into the Riflemargulisbacteria, primarily from sediments and groundwater, the Termititenax from insect gut microbiomes, and the Marinamargulisbacteria, from marine samples. We found that Marinamargulisbacteria amplicons were unusually distributed in size-fractionated samples from the sunlit photic and dark twilight zones of the ocean. Further, sequencing of wild marine protists rendered genomic information for distinct marinamargulisbacterial clades co-associated with uncultivated, non-photosynthetic Stramenopila and Opisthokonta protists. Phylogenomic analyses combining these data and available metagenome-assembled genomes (MAGs) and single-amplified genomes (SAGs) from sorted bacteria revealed new Marinamargulisbacteria lineages. The lineages delineate by their environment, forming clades comprising freshwater, marine pelagic, or sediment/hypoxic taxa. In conclusion, the remarkable diversity of Margulisbacteria indicates success in colonizing various habitats, potentially in a conserved strategy involving eukaryotic cells.

59 BASIC BIOLOGICAL SCIENCES↗

Stimulation of Dissimilatory Sulfate Reduction in Response to Sulfate in Microcosm Incubations from Two Contrasting Temperate Peatlands near Ithaca, NY, USA

Abstract Peatlands are responsible for over half of wetland methane emissions, yet major uncertainties remain regarding carbon flow, especially when increased availability of electron acceptors stimulate competing physiologies. We used microcosm incubations to study the effects of sulfate on microorganisms in two temperate peatlands, one bog and one fen. Three different electron donor treatments were used (13C-acetate, 13C-formate, and a mixture of 12C short-chain fatty acids) to elucidate the responses of sulfate-reducing bacteria (SRB) and methanogens to sulfate stimulation. Methane production was measured and metagenomic sequencing was performed, with only the heavy DNA fraction sequenced from treatments receiving 13C electron donors. Our data demonstrate stimulation of dissimilatory sulfate reduction in both sites, with contrasting community responses. In McLean Bog (MB), hydrogenotrophic Deltaproteobacteria and acetotrophic Peptococcaceae lineages of SRB were stimulated, as were lineages with unclassified dissimilatory sulfite reductases. In Michigan Hollow Fen (MHF), there was little stimulation of Peptococcaceae populations, and a small stimulation of Deltaproteobacteria SRB populations only in the presence of formate as electron donor. Sulfate stimulated an increase in relative abundance of reads for both oxidative and reductive sulfite reductases, suggesting stimulation of an internal sulfur cycle. Together, these data indicate a stimulation of SRB activity in response to sulfate in both sites, with a stronger growth response in MB than MHF. This study provides valuable insights into microbial community responses to sulfate in temperate peatlands and is an important first step to understanding how SRB and methanogens compete to regulate carbon flow in these systems.

Microbiology↗