Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “JGI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Mapping the soil microbiome functions shaping wetland methane emissions

Accounting for only 8% of Earth’s land cover, freshwater wetlands remain the foremost contributors to global methane emissions. Yet the microorganisms and processes underlying methane emissions from wetland soils remain poorly understood. Over a five-year period, we surveyed the microbial membership and in situ methane measurements from over 700 samples in one of the most prolific methane-emitting wetlands in the United States. We constructed a catalog of 2,502 metagenome-assembled genomes (MAGs), with more than half of the 70 bacterial and archaeal phyla sampled containing novel lineages. Integration of these data with 133 soil metatranscriptomes provided a genome-resolved view of the biogeochemical specialization and versatility expressed over wetland soil spatial and temporal gradients. Centimeter-scale depth differences best explained patterns of microbial community structure and transcribed functionalities, even more than land cover or temporal information. Moreover, while extended flooding restructured soil redox, this perturbation failed to reconfigure the transcriptional profiles of methane-cycling microorganisms, contrasting with theoretically expected responses to hydrological perturbations. Co-expression analyses, coupled with depth-resolved methane measurements, revealed the metabolisms and trophic structures most predictive of methane hotspots. Mapping the spatiotemporal transcriptional patterns on this compendium of biogeochemically classified soil-derived genomes begins to untangle the microbial carbon, energy, and nutrient processing contributing to wetland methane production.

MAG↗

ggtaxplot v 0.0.1

ggtaxplot is an R package designed to process and visualize taxonomic data through a taxonomic river plot. This package is ideal for researchers and data scientists who need to visualize taxonomic data. ggtaxplot function processes data and generates a taxonomic river plot, allowing users to visualize the distribution of taxa across different samples.

Coclet, Clement [Lawrence Berkeley National Labora↗

Global Archaeal Diversity Revealed Through Massive Data Integration: Uncovering Just Tip of Iceberg

The domain of Archaea has gathered significant interest for its ecological and biotechnological potential and its role in helping us to understand the evolutionary history of Eukaryotes. In comparison to the bacterial domain, the number of adequately described members in Archaea is relatively low, with less than 1000 species described. It is not clear whether this is solely due to the cultivation difficulty of its members or, indeed, the domain is characterized by evolutionary constraints that keep the number of species relatively low. Based on molecular evidence that bypasses the difficulties of formal cultivation and characterization, several novel clades have been proposed, enabling insights into their metabolism and physiology. Given the extent of global sampling and sequencing efforts, it is now possible and meaningful to question the magnitude of global archaeal diversity based on molecular evidence. To do so, we extracted all sequences classified as Archaea from 500 thousand amplicon samples available in public repositories. After processing through our highly conservative pipeline, we named this comprehensive resource the ‘Global Archaea Diversity’ (GAD), which encompassed nearly 3 million molecular species clusters at 97% similarity, and organized it into over 500 thousand genera and nearly 100 thousand families. Saline environments have contributed the most to the novel taxa of this previously unseen diversity. The majority of those 16S rRNA gene sequence fragments were verified by matches in metagenomic datasets from IMG/M. These findings reveal a vast and previously overlooked diversity within the Archaea, offering insights into their ecological roles and evolutionary importance while establishing a foundation for the future study and characterization of this intriguing domain of life.

59 BASIC BIOLOGICAL SCIENCES↗

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS↗

Hybridization breaks species barriers in long-term coevolution of a cyanobacterial population

Bacterial species often undergo rampant recombination yet maintain cohesive genomic identity. Ecological differences can generate recombination barriers between species and sustain genomic clusters in the short term. But can these forces prevent genomic mixing during long-term coevolution? Cyanobacteria in Yellowstone hot springs comprise several diverse species that have coevolved for hundreds of thousands of years, providing a rare natural experiment. By analyzing more than 300 single-cell genomes, we show that despite each species forming a distinct genomic cluster, much of the diversity within species is the result of hybridization driven by selection, which has mixed their ancestral genotypes. This widespread mixing is contrary to the prevailing view that ecological barriers can maintain cohesive bacterial species and highlights the importance of hybridization as a source of genomic diversity.

Evolutionary Biology↗

New approaches to secondary metabolite discovery from anaerobic gut microbes

The animal gut microbiome is a complex system of diverse, predominantly anaerobic microbiota with secondary metabolite potential. These metabolites likely play roles in shaping microbial community membership and influencing animal host health. As such, novel secondary metabolites from gut microbes hold significant biotechnological and therapeutic interest. Despite their potential, gut microbes are largely untapped for secondary metabolites, with gut fungi and obligate anaerobes being particularly under-explored. To advance understanding of these metabolites, culture-based and (meta)genome-based approaches are essential. Culture-based approaches enable isolation, cultivation, and direct study of gut microbes, and (meta)genome-based approaches utilize in silico tools to mine biosynthetic gene clusters (BGCs) from microbes that have not yet been successfully cultured. In this mini-review, we highlight recent innovations in this area, including anaerobic biofoundries like ExFAB, the NSF BioFoundry for Extreme & Exceptional Fungi, Archaea, and Bacteria. These facilities enable high-throughput workflows to study oxygen-sensitive microbes and biosynthetic machinery. Such recent advances promise to improve our understanding of the gut microbiome and its secondary metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

Xylose metabolic engineering of Issatchenkia orientalis for 3-hydroxypropionic acid production from cellulosic hydrolysate without nutrient supplementation

Bioconversion of lignocellulosic biomass offers a promising alternative to petroleum-based chemical production. However, inefficient xylose utilization and toxic compounds in cellulosic hydrolysate limit microbial fermentation, as the hydrolysate contains substantial amounts of xylose in addition to glucose. To address these challenges, we engineered Issatchenkia orientalis to produce 3-hydroxypropionic acid (3-HP) directly from sorghum hydrolysate under low-pH conditions. A heterologous xylose utilization pathway consisting of XYL1, XYL2, and XYL3 from Scheffersomyces stipitis was introduced into an engineered 3-HP producing strain, enabling efficient conversion of xylose to 3-HP. The engineered strain produced 46.8 g/L 3-HP from sorghum hydrolysate without nutrient supplementation. To eliminate the lag phase under low-pH conditions, fermentation was conducted at pH 6.0 for the first three days, after which pH control was discontinued and in situ 3-HP accumulation buffered the culture. This partial pH control strategy increased 3-HP productivity by 55% from 0.20 to 0.31 g/L∙h, while maintaining low-pH conditions. Introducing an additional copy of XYL2 further increased 3-HP titer to 53.5 g/L and the yield by 33%, from 0.30 to 0.40 g/g sugars, with pH reaching 4.5 at the end of fermentation. This represents one of the highest reported 3-HP titers and yields from cellulosic hydrolysate without additional nutrient supplementation. This work demonstrates a nutrient-independent and low-pH bioprocess for upgrading lignocellulosic hydrolysate into 3-HP, highlighting the industrial potential of engineered xylose-utilizing I. orientalis for sustainable production of platform chemicals from renewable feedstocks.

3-Hydroxypropionic acid↗

Origin of replication discovery for environmentally isolated Pantoea strain enables expression of heterologous proteins, pathways and products

Leveraging predicted origin sequences from a previously characterized groundwater plasmidome, we constructed a barcoded plasmid library to screen for previously unknown origins. Testing this library against a panel of representative bacterial strains led to the identification of 3 previously unknown origins that replicate in gram-negative bacteria not previously associated with these origin sequences. Experimental validation confirmed that a plasmid bearing origin 6911 as the sole origin could replicate with a copy number of 9 (±2) in Pantoea sp. MT58, a fast growing and metal tolerant, environmentally important bacterium. Plasmids based on this new origin were used to express the reporter protein GFP, and non-native metabolite pathways for the natural product indigoidine and the terpenoid compound isoprenol. Functional previously unknown origins of replication in such non-model organisms can expand the toolkit for genetic manipulations of both model and less-studied bacteria.

molecular biology↗

Advancing Protein Display on Bacterial Spores through an Extensive Survey of Coat Components

The profound stability of bacterial spores makes them a promising platform for biotechnological applications like biocatalysis, bioremediation, drug delivery, etc. However, though the Bacillus subtilis spore is composed of >40 types of proteins, only ∼12 have been explored as fusion carriers for protein display. Here, we assessed the suitability of 33 spore proteins (SPs) as enzyme display carriers by direct allele tagging at native genomic loci. Of the 33 SPs investigated, 26 formed functional fusions with β-glucuronidase (GUS)─a ∼272 kDa homotetramer. This almost triples the number of SPs assessed for enzyme display and doubles the number of functional fusions documented in the literature. We quantitatively assessed 1) SP promoter activation dynamics, 2) GUS activity on spores, 3) surface availability, and 4) protection from thermal and proteolytic degradation. Multicopy expression and pairwise coexpression of the most promising SP-GUS fusions highlighted the complexity of spore structure/assembly and the difficulty in predicting compatibility between different SP fusions. We also assessed the suitability of engineered spores to degrade PET (polyethylene terephthalate) films and found that surface-exposed SPs were most effective. Beyond the broad survey, a key outcome of our work was the identification of SscA (small spore coat assembly protein A) as an effective spore display carrier. SscA supported enzyme activity at least 4-fold higher than any other SP, including the well-established anchor, CotY. We attribute this to its promoter, which demonstrated early and sustained activation relative to other SPs and its small size (∼3 kDa), which likely minimally interferes with enzyme folding, oligomerization, and activity. Labeling and genetic studies, its hydrophobic nature, and low surface availability suggest that SscA assembles within the inner spore coat, which makes it stabilizing and suitable for many biocatalytic applications. Overall, this work serves as a knowledge base to advance the biotechnological utility of B. subtilis spores.

Bacillus subtilis↗

Continental-scale integration of soil metagenomes and organic matter chemistry reveals ubiquitous microbial capacity for chemically-recalcitrant carbon decomposition

Soil organic matter (SOM) decomposition by microorganisms is a major uncertainty in predicting terrestrial carbon–atmosphere feedbacks, partly because we lack understanding of the microbial diversity involved in depolymerizing different carbon pools across environmental gradients. We address this gap using a continental-scale dataset pairing shotgun metagenomes with high-resolution SOM chemistry, assembling 0.76 Tbp of prokaryotic MAGs (828 genomes) and identifying 66,727 SOM molecules from 47 standardized U.S. soil cores selected using respiration rates from 106 soils. Integrating these datasets reveals widespread microbial potential for depolymerizing chemically-recalcitrant SOM previously considered stable. We uncover complementary metabolic specialization between genera affiliated with two abundant bacterial orders, Rhizobiales and Chthoniobacterales, and an archaeal order, Nitrososphaerales. This metabolic partitioning is consistent across soil depths and activity levels, suggesting coordinated decomposition of complex SOM through distinct but complementary biochemical strategies. The metabolic potential for depolymerization of chemically-recalcitrant compounds is supported by the abundance of these molecules across the soils, as indicated by Fourier-Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS), and by flux balance analysis of metabolic models. Our results show that a substantial portion of ostensibly stable SOM remains vulnerable to microbial decomposition, a mechanism not captured in current Earth System Models.

Song, Young C. [Pacific Northwest National Laborat↗

Two decades of bacterial ecology and evolution in a freshwater lake

Ecology and evolution are considered distinct processes that interact on contemporary time scales in microbiomes. Here, to observe these processes in a natural system, we collected a two-decade, 471-metagenome time series from Lake Mendota (Wisconsin, USA). We assembled 2,855 species-representative genomes and found that genomic change was common and frequent. By tracking strain composition via single nucleotide variants, we identified cyclical seasonal patterns in 80% and decadal shifts in 20% of species. In the dominant freshwater family Nanopelagicaceae, environmental extremes coincided with shifts in strain composition and positive selection of amino acid and nucleic acid metabolism genes. Further, these genes identify organic nitrogen compounds as potential drivers of freshwater responses to global change. Seasonal and long-term strain dynamics could be regarded as ecological processes or, equivalently, as evolutionary change. Rather than as distinct interacting processes, we propose a conceptualization of ecology and evolution as a continuum to better describe change in microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

Phenogenomics reveals the ecology and evolution of Trichoderma fungi for sustainable agriculture

Trichoderma fungi support sustainable agriculture by suppressing plant diseases and improving crop performance. However, emerging pathogenicity of Trichoderma warrants further ecological and genetic characterization. Here we used machine learning to correlate genomic data from 37 Trichoderma strains with over 140 phenotypic traits, spanning metabolic versatility, biotic interactions, stress tolerance and reproductive strategies. We determined Trichoderma to be an ancient, genetically cohesive and physiologically diverse genus with spores capable of germination in water and dispersal via air and water droplets. Metabolic preferences indicate universal adaptation to mycoparasitism and to niches like arboreal microbial mats, alongside broader saprotrophic versatility. Our analyses are consistent with character displacement among close relatives and convergent evolution in distant lineages, with both processes shaping ecological plasticity and traits including dispersal modes, terrestrialization or endophytism. Our findings reveal that while some Trichoderma species show traits of biosafety concern, its vast ecophysiological diversity enables the development of safe, targeted bioeffectors.

Steindorff, Andrei S. [USDOE Joint Genome Institut↗

An expanded registry of candidate cis -regulatory elements

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq, massively parallel reporter assay, CRISPR perturbation and transgenic mouse assays have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Moore, Jill E. [Univ. of Massachusetts, Worchester↗

Multiomics and deep learning dissect regulatory syntax in human development

Transcription factors establish cell identity during development by binding regulatory DNA in a sequence-specific manner, often promoting local chromatin accessibility and regulating gene expression1. Mapping accessible chromatin offers critical insights into transcriptional control, but available datasets for human development are restricted to bulk tissue, single organs or single modalities2. Here we present the Human Development Multiomic Atlas, a single-cell atlas of chromatin accessibility and gene expression from 817,740 fetal cells across 12 organs, spanning 203 cell types and more than 1 million candidate cis-regulatory elements, many of which exhibit organ-specific in vivo enhancer activity. Deep learning models trained to predict accessibility from local DNA sequence unravel a comprehensive lexicon of motifs that influence accessibility, including composite motifs exhibiting distinct syntactic constraints that are predicted to mediate transcription factor cooperativity. We identify ‘hard’ syntactic rules requiring precise motif spacing and orientation, ‘soft’ rules allowing flexible motif arrangements, and ubiquitous motifs inhibiting accessibility. Model-based interpretation of genetic variants reveals that disruption of motifs with positive and negative effects is associated with concordant effects on gene expression. Our work delineates how motif syntax governs cell-type-specific chromatin accessibility and provides a foundational resource for decoding cis-regulatory logic and interpreting genetic variation during human development.

59 BASIC BIOLOGICAL SCIENCES↗

Cell-free synthetic biology for natural product biosynthesis and discovery

Natural products have applications as biopharmaceuticals, agrochemicals, and other high-value chemicals. However, there are challenges in isolating natural products from their native producers (e.g. bacteria, fungi, plants). In many cases, synthetic chemistry or heterologous expression must be used to access these important molecules. The biosynthetic machinery to generate these compounds is found within biosynthetic gene clusters, primarily consisting of the enzymes that biosynthesise a range of natural product classes (including, but not limited to ribosomal and nonribosomal peptides, polyketides, and terpenoids). Cell-free synthetic biology has emerged in recent years as a bottom-up technology applied towards both prototyping pathways and producing molecules. Recently, it has been applied to natural products, both to characterise biosynthetic pathways and produce new metabolites. This review discusses the core biochemistry of cell-free synthetic biology applied to metabolite production and critiques its advantages and disadvantages compared to whole cell and/or chemical production routes. Specifically, we review the advances in cell-free biosynthesis of ribosomal peptides, analyse the rapid prototyping of natural product biosynthetic enzymes and pathways, highlight advances in novel antimicrobial discovery, and discuss the rising use of cell-free technologies in industrial biotechnology and synthetic biology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

A haplotype-resolved, chromosome-scale genome assembly for the southern live oak, Quercus virginiana

Hybridization is a major force driving diversification, migration, and adaptation in Quercus species. While population genetics and phylogenetics have traditionally been used for studying these processes, advances in sequencing technology now enable us to incorporate comparative and pan-genomic approaches as well. Here, we present a highly contiguous, chromosome-scale and haplotype-resolved genome assembly for the southern live oak, Quercus virginiana, the first reference genome for section Virentes, as part of the American Campus Tree Genomes program. Originating from a clone of Auburn University's historic “Toomer's Oak,” this assembly contributes to the pool of genomic resources for investigating recombination, haplotype variation, and structural genomic changes influencing hybridization potential in this clade and across Quercus. It also provides insights into the architecture of the putative centromeric regions within the genus. Alongside other oak references, the Q. virginiana genome will support research into the evolution and adaptation of the Quercus genus.

Quercus virginiana↗