Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Heritable gene editing in tomato through viral delivery of isopentenyl transferase and single-guide RNAs to latent axillary meristematic cells

Realizing the full potential of genome editing for crop improvement has been slow due to inefficient methods for reagent delivery and the reliance on tissue culture for creating gene-edited plants. RNA viral vectors offer an alternative approach for delivering gene engineering reagents and bypassing the tissue culture requirement. Viruses, however, are often excluded from the shoot apical meristem, making virus-mediated gene editing inefficient in some species. Here, we developed effective approaches for generating gene-edited shoots in Cas9-expressing transgenic tomato plants using RNA virus-mediated delivery of single-guide RNAs (sgRNAs). RNA viral vectors expressing sgRNAs were either delivered to leaves or sites near axillary meristems. Trimming of the apical and axillary meristems induced new shoots to form from edited somatic cells. To further encourage the induction of shoots, we used RNA viral vectors to deliver sgRNAs along with the cytokinin biosynthesis gene, isopentenyl transferase. Abundant, phenotypically normal, gene-edited shoots were induced per infected plant with single and multiplexed gene edits fixed in the germline. The use of viruses to deliver both gene editing reagents and developmental regulators overcomes the bottleneck in applying virus-induced gene editing to dicotyledonous crops such as tomato and reduces the dependency on tissue culture.

59 BASIC BIOLOGICAL SCIENCES↗

Targeted seed EMS mutagenesis reveals a basic helix–loop–helix transcription factor underlying male sterility in sorghum

Abstract Forward genetic screens of mutant populations are fundamental for functional genomics studies. However, isolating independent mutant alleles to molecularly identify causal genes is challenging in species recalcitrant to genetic manipulation. Here, we demonstrate that classic seed ethyl methanesulfonate (EMS) mutagenesis coupled with genome sequencing can overcome this limitation in sorghum. We used this method to generate new mutant alleles of sorghum MALE STERILE 8 (MS8) and identified the causal locus for the ms8 phenotype as Sobic.004G270900, which encodes the sorghum ortholog of maize bhlh122, a basic helix–loop–helix (bHLH) transcription factor required for male fertility in maize. Bulked segregant analysis mapped ms8-1 to a region on chromosome 4 containing Sobic.004G270900. Seeds from heterozygous MS8/ms8-1 plants were mutagenized and screened for chimeric inflorescences containing sectors with white, sterile anthers resembling the ms8-1 homozygous phenotype. DNA sequencing of sterile and fertile sectors from a single chimeric inflorescence revealed two mutations in Sobic.004G270900 within the sterile sector, but not the fertile sector. Isolation of this loss-of-function allele (ms8-2) established Sobic.004G270900 as the causative locus for male sterility in the ms8 mutant. We generated additional alleles of MS8 in a different genetic background using CRISPR/Cas9-based gene editing, where deletions in Sobic.004G270900 also resulted in male sterility. Our work identified a gene underlying male sterility in sorghum and provides a novel and straightforward genetic tool for researchers who lack access to advanced transformation facilities to validate gene candidates. Unlike gene editing, no prior knowledge of candidate genes is required for targeted seed EMS mutagenesis to aid identification of causal loci.

Genetics & Heredity↗

Nanoporous Materials Genome Center Final Technical Report

Nanoporous materials (NPMs), including zeolites/zeotypes, metal-organic frameworks (MOFs), covalent organic frameworks, polymers with intrinsic microporosity, and molecular cages, possess enormous potential in diverse areas relevant to the DOE Office of Science Basic Energy Sciences (BES) mission and objectives. The Nanoporous Materials Genome Center (NMGC) has developed exascale-ready software, computational/theoretical chemistry methods, and data-driven science approaches that enable (i) the de-novo design of functional NPMs for chemical separation and catalysis tasks of increasing complexity, (ii) the discovery of the most promising functional NPMs from databases of synthesized and hypothetical adsorbent structures and the optimization of process conditions for specific applications, and (iii) the microscopic-level understanding of the fundamental interactions underlying the function of NPMs including hierarchical architectures, composite materials, responsive frameworks that may undergo phase transitions or post-synthetic modifications, and materials containing defects, partial disorder, or interfaces. A pivotal part of the NMGC project has been a tight collaboration between leading experimental groups for synthesis and characterization of NPMs and of computational groups that allowed for iterative feedback. The NMGC project has resulted in the publication of more than 290 research and review articles including more than 60 publications in high-impact journals and more than 15 journal covers. NMGC publications have already received more than 20,000 citations (with more than 3,000 citations per year in 2021, 2022, and 2023) and contribute to an h-index of more than 72. The NMGC award has supported collaborative research involving 28 research groups and contributed to the training of more than 40 postdocs, more than 60 graduate students, and more than 20 undergraduate students with broad expertise in data-driven science approaches, computational chemistry methods, and high-performance computing, in addition to the skills to thrive in an integrated experimental and computational research environment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CRAGE-RB-PI-seq reveals transcriptional dynamics of plant-associated bacteria during root colonization

Plant roots release a wide array of metabolites into the rhizosphere, shaping microbial communities and their functions. While metagenomics has expanded our understanding of these communities, little is known about the physiology of their members in host environments. Transcriptome analysis via RNA sequencing is a common approach to learning more, but its use has been challenging because of low bacterial biomass and interference from plant RNA. To overcome this, we developed a randomly-barcoded promoter-library insertion sequencing (RB-PI-seq) combined with chassis-independent recombinase-assisted genome engineering (CRAGE). Using Pseudomonas simiae WCS417 as a model rhizobacterium, this method enabled targeted amplification of barcoded transcripts, bypassing plant RNA interference and allowing measurement of thousands of promoter activities during Arabidopsis root colonization. Our analysis revealed temporally resolved transcriptional regulation, including those associated with cell growth, chemotaxis, plant immune suppression, biofilm formation, and stress responses, reflecting the coordinated physiological adaptation to the root environment. Additionally, we discovered that transcriptional activation of xanthine dehydrogenase and a lysozyme inhibitor is crucial for evading plant immune systems. This framework is scalable to other bacterial species and provides new opportunities for understanding rhizobacterial gene regulation in native environments.

59 BASIC BIOLOGICAL SCIENCES↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Genome‐enabled exploration of microbial ecology and evolution in the sea: a rising tide lifts all boats

Summary As a young bacteriologist just launching my career during the early days of the ‘microbial revolution’ in the 1980s, I was fortunate to participate in some early discoveries, and collaborate in the development of cross‐disciplinary methods now commonly referred to as "metagenomics". My early scientific career focused on applying phylogenetic and genomic approaches to characterize ‘wild’ bacteria, archaea and viruses in their natural habitats, with an emphasis on marine systems. These central interests have not changed very much for me over the past three decades, but knowledge, methodological advances and new theoretical perspectives about the microbial world certainly have. In this invited ‘How we did it’ perspective, I trace some of the trajectories of my lab's collective efforts over the years, including phylogenetic surveys of microbial assemblages in marine plankton and sediments, development of microbial community gene‐ and genome‐enabled surveys, and application of genome‐guided, cultivation‐independent functional characterization of novel enzymes, pathways and their relationships to in situ biogeochemistry. Throughout this short review, I attempt to acknowledge, all the mentors, students, postdocs and collaborators who enabled this research. Inevitably, a brief autobiographical review like this cannot be fully comprehensive, so sincere apologies to any of my great colleagues who are not explicitly mentioned herein. I salute you all as well!

59 BASIC BIOLOGICAL SCIENCES↗

Nuclear and chloroplast genome engineering of a productive non-model alga Desmodesmus armatus: Insights into unusual and selective acquisition mechanisms for foreign DNA

Despite the tremendous potential of algae to contribute to a future bioeconomy, there are practical and theoretical limitations to how well naturally sourced species and strains can perform in an outdoor setting. The application of biotechnology to modulate and engineer algae metabolism or to increase performance, resilience, or produce novel compounds, offers opportunities to overcome some of the major commercialization barriers. There are numerous approaches reported in the literature having variable success on genetic engineering of algae with non-model algae often presenting unique challenges to genetic engineering. We report here on successful nuclear and chloroplast genomic integration of selection marker resistance in the non-model alga Desmodesmus armatus. Nuclear transformation was accomplished using both electroporation and Agrobacterium-mediated approaches. However, in all surviving transformants, DNA integration was accompanied by excision and/or rearrangement of the gene of interest and fluorescence reporter coding sequences. Similarly, chloroplast transformation was successfully accomplished using a biolistic DNA delivery method. For these transformants, we also observed off-target mutations in the chloroplast genome, not previously observed in other, more routinely used, algae species. Finally, we present insights into potential mechanisms for these observed truncations, rearrangements, and mutations in D. armatus.

59 BASIC BIOLOGICAL SCIENCES↗

Complementary Metagenomic Approaches Improve Reconstruction of Microbial Diversity in a Forest Soil

ABSTRACT Soil ecosystems harbor diverse microorganisms and yet remain only partially characterized as neither single-cell sequencing nor whole-community sequencing offers a complete picture of these complex communities. Thus, the genetic and metabolic potential of this “uncultivated majority” remains underexplored. To address these challenges, we applied a pooled-cell-sorting-based mini-metagenomics approach and compared the results to bulk metagenomics. Informatic binning of these data produced 200 mini-metagenome assembled genomes (sorted-MAGs) and 29 bulk metagenome assembled genomes (MAGs). The sorted and bulk MAGs increased the known phylogenetic diversity of soil taxa by 7.2% with respect to the Joint Genome Institute IMG/M database and showed clade-specific sequence recruitment patterns across diverse terrestrial soil metagenomes. Additionally, sorted-MAGs expanded the rare biosphere not captured through MAGs from bulk sequences, exemplified through phylogenetic and functional analyses of members of the phylum Bacteroidetes . Analysis of 67 Bacteroidetes sorted-MAGs showed conserved patterns of carbon metabolism across four clades. These results indicate that mini-metagenomics enables genome-resolved investigation of predicted metabolism and demonstrates the utility of combining metagenomics methods to tap into the diversity of heterogeneous microbial assemblages. IMPORTANCE Microbial ecologists have historically used cultivation-based approaches as well as amplicon sequencing and shotgun metagenomics to characterize microbial diversity in soil. However, challenges persist in the study of microbial diversity, including the recalcitrance of the majority of microorganisms to laboratory cultivation and limited sequence assembly from highly complex samples. The uncultivated majority thus remains a reservoir of untapped genetic diversity. To address some of the challenges associated with bulk metagenomics as well as low throughput of single-cell genomics, we applied flow cytometry-enabled mini-metagenomics to capture expanded microbial diversity from forest soil and compare it to soil bulk metagenomics. Our resulting data from this pooled-cell sorting approach combined with bulk metagenomics revealed increased phylogenetic diversity through novel soil taxa and rare biosphere members. In-depth analysis of genomes within the highly represented Bacteroidetes phylum provided insights into conserved and clade-specific patterns of carbon metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

The GREEN ‘omics of Nutrient Feedbacks to Soil Warming

The GREEN ‘omics of Nutrient Feedbacks in Soil project advanced the DOE Biological and Environmental Research (BER) mission by developing and applying isotope-enabled ’omics tools to understand how soil microbes regulate carbon and nutrient cycling. Guided by the Growth Rate, growth Efficiency, and stoichiometry of Essential Nutrients (GREEN ’omics) framework, the project aimed to build a predictive, systems-level understanding of microbial traits that control ecosystem biogeochemistry. In a collaboration among Northern Arizona University (lead), West Virginia University, Lawrence Livermore National Laboratory, and Pacific Northwest National Laboratory, we combined quantitative stable isotope probing (qSIP), Chip-SIP, NanoSIMS, and genome-resolved metagenomics across long-term experiments in Arctic, boreal, temperate, and tropical ecosystems. The project produced three key outcomes: 1) We showed that community-weighted temperature sensitivities of bacterial growth (Q10) can predict ecosystem-scale soil respiration responses across diverse soils. 2) We provided the first in situ evidence for density-dependent population dynamics in soil bacteria and demonstrated that nutrient additions intensify competition, concentrating carbon use into fewer taxa. 3) We improved and extended isotope-enabled ’omics methods by quantifying qSIP measurement error to guide experimental design and coupling SIP with genome-resolved metagenomics to reveal cross-kingdom interactions among bacteria, fungi, and viruses. Together, these results show that a small number of microbial traits and taxa exert disproportionate control over soil carbon and nutrient cycling, providing critical data and methods to improve representation of microbial processes in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

JGI Plant Transformation Workshop, May 20-21, 2025

Domestic biomass crops such as sorghum, switchgrass, Miscanthus, and poplar can provide United States industries with renewable feedstocks while also supporting low-input farming systems and strengthening supply chains for biofuels, biochemicals and biomaterials. The U.S. leads globally in biomass crop genomics, yet progress in engineering traits is constrained by slow, genotype-dependent transformation methods and lengthy Design-Build-Test-Learn (DBTL) cycles. At a May 2025 workshop, a panel of experts recommended establishing a DOE Plant Transformation Capability (PTC) to overcome these barriers. The PTC would unite two missions: advancing research to achieve genotype-independent, automated methods, and delivering scalable transformation services through a user-facility model. With expected gains of 10–100x in efficiency, including transformation and cost reduction, the PTC would accelerate the path from discovery to engineered plants, expand community access and training, and support downstream applications and workflows including field trials and regulatory navigation. By enabling rapid and predictable crop engineering, the PTC would strengthen U.S. supply chains, enhance industrial competitiveness, and ensure that DOE’s genomic investments deliver national impact.

09 BIOMASS FUELS↗

Enumeration as a Tool for Structure Solution: A Materials Genomic Approach to Solving the Cation-Ordered Structure of Na 3 V 2 (PO 4 ) 2 F 3

While powder diffraction methods are routinely utilized to optimize structural models for compounds whose crystal structures are known, the determination of unknown structures is far more challenging. When the unknown structure is large, structure solution can become a virtually intractable problem using standard structure solution methodologies, especially when the space group cannot be unambiguously resolved. One such system is the promising Na-ion battery cathode material Na 3 V 2 (PO 4 ) 2 F 3 whose high temperature and room temperature structures were previously solved, but whose more complex low-temperature structure could not be determined. Here, a novel materials genomic approach is demonstrated for the solution of the unknown 100 K structure of Na 3 V 2 (PO 4 ) 2 F 3 in which enumeration methods are first used to generate a large number (~3,000) of trial structures based on plausible orderings of Na ions and then automated Rietveld refinements are carried out to optimize each of these trial structures. Based on both the analysis of the ensemble of optimized trial structures and the density functional theory energy minimization of selected trial structures, the 100 K structure of Na 3 V 2 (PO 4 ) 2 F 3 is best described as belonging to the space group A2 1 am with unit cell dimensions of a = 9.01928(4), b = 27.1379(1), c = 10.73307(5). The 100 K unit cell has a large volume of 2627.07(2) Å 3 with Z = 12 and 33 independent crystallographic sites (9 Na, 3 V, 3 P, 12 O, and 6 F) that is 3x and 6x larger than the room- and high-temperature polymorphs of this phase, respectively. Finally, the novel methods described here will be generally applicable for the solution of the complex cation-ordered structures that commonly occur for battery materials.

36 MATERIALS SCIENCE↗

Mycoparasites, Gut Dwellers, and Saprotrophs: Phylogenomic Reconstructions and Comparative Analyses of Kickxellomycotina Fungi

Improved sequencing technologies have profoundly altered global views of fungal diversity and evolution. High-throughput sequencing methods are critical for studying fungi due to the cryptic, symbiotic nature of many species, particularly those that are difficult to culture. However, the low coverage genome sequencing (LCGS) approach to phylogenomic inference has not been widely applied to fungi. Here we analyzed 171 Kickxellomycotina fungi using LCGS methods to obtain hundreds of marker genes for robust phylogenomic reconstruction. Additionally, we mined our LCGS data for a set of nine rDNA and protein coding genes to enable analyses across species for which no LCGS data were obtained. The main goals of this study were to: 1) evaluate the quality and utility of LCGS data for both phylogenetic reconstruction and functional annotation, 2) test relationships among clades of Kickxellomycotina, and 3) perform comparative functional analyses between clades to gain insight into putative trophic modes. In opposition to previous studies, our nine-gene analyses support two clades of arthropod gut dwelling species and suggest a possible single evolutionary event leading to this symbiotic lifestyle. Furthermore, we resolve the mycoparasitic Dimargaritales as the earliest diverging clade in the subphylum and find four major clades of Coemansia species. Finally, functional analyses illustrate clear variation in predicted carbohydrate active enzymes and secondary metabolites (SM) based on ecology, that is biotroph versus saprotroph. Saprotrophic Kickxellales broadly lack many known pectinase families compared with saprotrophic Mucoromycota and are depauperate for SM but have similar numbers of predicted chitinases as mycoparasitic.

59 BASIC BIOLOGICAL SCIENCES↗

Virus-induced gene editing free from tissue culture

Virus-induced gene editing (VIGE) has reached an inflection point. Although conceived as an alternative to traditional methods of producing gene-edited plants, VIGE has historically relied on the very technologies it was meant to supersede—specifically, tissue-culture-mediated transgenesis. Recent VIGE innovations, however, have finally proved its viability as an independent method for plant gene editing. Here we discuss the advances in plant genome engineering VIGE may unlock, what progress has been made towards achieving these advances and the challenges that continue to impede that progress.

54 ENVIRONMENTAL SCIENCES↗

Decomposing a San Francisco estuary microbiome using long-read metagenomics reveals species- and strain-level dominance from picoeukaryotes to viruses

ABSTRACT Although long-read sequencing has enabled obtaining high-quality and complete genomes from metagenomes, many challenges still remain to completely decompose a metagenome into its constituent prokaryotic and viral genomes. This study focuses on decomposing an estuarine metagenome to obtain a more accurate estimate of microbial diversity. To achieve this, we developed a new bead-based DNA extraction method, a novel bin refinement method, and obtained 150 Gbp of Nanopore sequencing. We estimate that there are ~500 bacterial and archaeal species in our sample and obtained 68 high-quality bins (>90% complete, <5% contamination, ≤5 contigs, contig length of >100 kbp, and all ribosomal and tRNA genes). We also obtained many contigs of picoeukaryotes, environmental DNA of larger eukaryotes such as mammals, and complete mitochondrial and chloroplast genomes and detected ~40,000 viral populations. Our analysis indicates that there are only a few strains that comprise most of the species abundances. IMPORTANCE Ocean and estuarine microbiomes play critical roles in global element cycling and ecosystem function. Despite the importance of these microbial communities, many species still have not been cultured in the lab. Environmental sequencing is the primary way the function and population dynamics of these communities can be studied. Long-read sequencing provides an avenue to overcome limitations of short-read technologies to obtain complete microbial genomes but comes with its own technical challenges, such as needed sequencing depth and obtaining high-quality DNA. We present here new sampling and bioinformatics methods to attempt decomposing an estuarine microbiome into its constituent genomes. Our results suggest there are only a few strains that comprise most of the species abundances from viruses to picoeukaryotes, and to fully decompose a metagenome of this diversity requires 1 Tbp of long-read sequencing. We anticipate that as long-read sequencing technologies continue to improve, less sequencing will be needed.

Lui, Lauren M.↗

Growing field of materials informatics: databases and artificial intelligence

The paradigm of molecular discovery in the chemical and pharmaceutical industry has followed a repetitive succession of screening and synthesis, involving the analysis of individual molecules that were both natural and produced. This ability to generate and screen libraries of compounds has found an echo in solid-state physics with the demand to explore and produce new materials for testing. In response to this demand, a golden age of materials discovery is being developed, with progress on important areas of both basic science and device applications. The confluence of theoretical and simulation methods, together with the availability of computation resources, has established the “materials genome” approach that is used by a growing number of research groups around the world with the goal of innovating on materials through systematic discovery. Here, an overview of this group of methodologies in tackling the ever-increasing complexity of computational materials science simulations is provided. Computational simulation is highlighted as a major component of rational design and synthesis of new materials with targeted properties, describing progress on databases and large data treatment. Tools for new materials discovery, including progress on the deployment of new data repositories, the implementation of high-throughput simulation approaches, and the development of artificial intelligence algorithms, are discussed.

36 MATERIALS SCIENCE↗

Inter-Kingdom Viral Interactions

Please cite as : Josué A. Rodríguez-Ramos, Amy E. Zimmerman, Ruonan Wu, Sheryl Bell, Trinidad Alfaro, Kirsten Hofmockel, William C. Nelson. 2025. Inter-Kingdom Viral Interactions. [Data Set] PNNL DataHub. This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the above citations for the data package and associated manuscript. Deciphering viral ecology in soils is challenging due to their high physiochemical and community complexity. To enhance detection of sub-communities of DNA and RNA viruses, we applied fractionation approaches to soils collected across a moisture gradient from a grassland field experiment. Analyses included metagenomics and metatranscriptomics of size-fractionated extracellular viruses (i.e., DNA and RNA viromes), metagenomics of bacteria/archaea- or eukaryote-enriched samples, and whole soil metatranscriptomes with rRNA-depletion or polyadenylation enrichment. While RNA virome and whole soil RNA methods captured similar viral diversity, RNA viromes identified longer, higher-quality genomes. Further, we showed that significantly more DNA viruses were active in higher moisture than lower moisture samples, whereas responses by overall diversity vary by genome type (DNA versus RNA genomes). Finally, we demonstrate the power of fractionation approaches for identifying distinct viral communities that infect unique hosts, which has significant implications for ecological investigations, particularly related to interkingdom interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Veridical data science

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, composed of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. As part of the PCS workflow, we develop PCS inference procedures, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others. Moreover, we demonstrate its favorable performance over existing methods in terms of receiver operating characteristic (ROC) curves in high-dimensional, sparse linear model simulations, including a wide range of misspecified models. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

97 MATHEMATICS AND COMPUTING↗

A standardized quantitative analysis strategy for stable isotope probing metagenomics

ABSTRACT Stable isotope probing (SIP) facilitates culture-independent identification of active microbial populations within complex ecosystems through isotopic enrichment of nucleic acids. Many DNA-SIP studies rely on 16S rRNA gene sequences to identify active taxa, but connecting these sequences to specific bacterial genomes is often challenging. Here, we describe a standardized laboratory and analysis framework to quantify isotopic enrichment on a per-genome basis using shotgun metagenomics instead of 16S rRNA gene sequencing. To develop this framework, we explored various sample processing and analysis approaches using a designed microbiome where the identity of labeled genomes and their level of isotopic enrichment were experimentally controlled. With this ground truth dataset, we empirically assessed the accuracy of different analytical models for identifying active taxa and examined how sequencing depth impacts the detection of isotopically labeled genomes. We also demonstrate that using synthetic DNA internal standards to measure absolute genome abundances in SIP density fractions improves estimates of isotopic enrichment. In addition, our study illustrates the utility of internal standards to reveal anomalies in sample handling that could negatively impact SIP metagenomic analyses if left undetected. Finally, we present SIPmg , an R package to facilitate the estimation of absolute abundances and perform statistical analyses for identifying labeled genomes within SIP metagenomic data. This experimentally validated analysis framework strengthens the foundation of DNA-SIP metagenomics as a tool for accurately measuring the in situ activity of environmental microbial populations and assessing their genomic potential. IMPORTANCE Answering the questions, “who is eating what?” and “who is active?” within complex microbial communities is paramount for our ability to model, predict, and modulate microbiomes for improved human and planetary health. These questions can be pursued using stable isotope probing to track the incorporation of labeled compounds into cellular DNA during microbial growth. However, with traditional stable isotope methods, it is challenging to establish links between an active microorganism’s taxonomic identity and genome composition while providing quantitative estimates of the microorganism’s isotope incorporation rate. Here, we report an experimental and analytical workflow that lays the foundation for improved detection of metabolically active microorganisms and better quantitative estimates of genome-resolved isotope incorporation, which can be used to further refine ecosystem-scale models for carbon and nutrient fluxes within microbiomes.

54 ENVIRONMENTAL SCIENCES↗