Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Joint Genome Institute”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Genome-scale Design and Engineering of Non-model Yeast Organisms for Production of Biofuels and Bioproducts

The overall goal of this project was to develop genome-scale design and engineering tools for two non-model yeast organisms including Rhodotorula toruloides and Issatchenkia orientalis to produce high-levels of fatty acids-derived products and organic acids, respectively. The project was performed between 9/15/2017 and 9/14/2024 (the last two-years were no-cost extensions). The team consisted of Huimin Zhao (Lead PI) and Christopher Rao (Co-PI) from the University of Illinois at Urbana-Champaign (UIUC), Costas Maranas (Co-PI) from the Pennsylvania State University, Joshua Rabinowitz (Co-PI) and Martin Wuhr (Co-PI) from Princeton University, and Yasuo Yoshikuni (Co-PI) from the DOE Joint Genome Institute. The team has made great progress in both tool development and fundamental understanding of these two non-model yeasts. In total, there were 40 research publications (one of them is still under review) and one patent application as well as numerous oral presentations.

60 APPLIED LIFE SCIENCES↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

1000 Soils Pilot Dataset, version 8, May 2025

This record hosts data generated by the 1000 Soils Pilot. Data will be updated as more become available. Please see the most recent data upload for current data. A beta visualization tool is available for some data types at https://shinyproxy.emsl.pnnl.gov/app/1000soils. Please submit any suggestions or comments through the 'contact' tab. We are actively working to improve visualizations and value all feedback. Data completed include: Geochemistry, texture, respiration, and enzyme activities FTICR-MS organic matter chemistry Microbial biomass C and N TOC/TDN of water-extractable OM X-ray computed tomography (derived metrics available here, raw data available upon request) Metagenomes; a variety of data formats are available upon request Soil hydraulic properties Data in progress: LC-MS/MS in development, timeline TBD, inquire for status 1000S_processed_BGC_summary.csv contains all available biogeochemical data; microbial biomass C and N; and TOC/TDN of water-extractable OM; and 1000S_Tomography.xslx contains a summary of data generated via X-ray computed tomography. icr_v2_corems2.csv contains FTICR-MS data processed by CoreMS version 2. These data are merged by formula across instrument runs to enable cross-sample comparisons. Technical replicates are merged by retaining peaks present in 2 out of 3 replicates. 1000Soils_Metadata_Site_Mastersheet_v1.csv contains site information. Soil Hydraulics_corrected_02042025.xlsx contains soil hydraulics information. Readme File_v4.xlsx is the readme file. Please contact the MONet project (monet.emsl@pnnl.gov) or Emily Graham (emily.graham@pnnl.gov) with questions. The following file and all raw data are available upon request: icr_by_mass_for_single_sample_analysis_only.csv contains FTICR-MS data processed by CoreMS and is intended for usage in the calculation of biochemical transformations within samples only. These data are not acceptable for cross-sample comparison of masses because they are from multiple instrument runs. For more information, please see: https://www.emsl.pnnl.gov/monet and https://sc-data.emsl.pnnl.gov/monet Acknowledgment: Soil data were provided by the Molecular Observation Network (MONet) at the Environmental Molecular Sciences Laboratory (https://ror.org/04rc0xn13), a DOE Office of Science user facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. The work (proposal: 10.46936/10.25585/60008970) conducted by the U.S. Department of Energy, Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science user facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231. The Molecular Observation Network (MONet) database is an open, FAIR, and publicly available compilation of the molecular and microstructural properties of soil. Data in the MONet open science database can be found at https://sc-data.emsl.pnnl.gov/.

biogeochemistry↗

Comparative genomics of Aspergillus nidulans and section Nidulantes

Aspergillus nidulans is an important model organism for eukaryotic biology and the reference for the section Nidulantes in comparative studies. In this study, we de novo sequenced the genomes of 25 species of this section. Whole-genome phylogeny of 34 Aspergillus species and Penicillium chrysogenum clarifies the position of clades inside section Nidulantes. Comparative genomics reveals a high genetic diversity between species with 684 up to 2433 unique protein families. Furthermore, we categorized 2118 secondary metabolite gene clusters (SMGC) into 603 families across Aspergilli, with at least 40 % of the families shared between Nidulantes species. Genetic dereplication of SMGC and subsequent synteny analysis provides evidence for horizontal gene transfer of a SMGC. Proteins that have been investigated in A. nidulans as well as its SMGC families are generally present in the section Nidulantes, supporting its role as model organism. The set of genes encoding plant biomass-related CAZymes is highly conserved in section Nidulantes, while there is remarkable diversity of organization of MAT-loci both within and between the different clades. This study provides a deeper understanding of the genomic conservation and diversity of this section and supports the position of A. nidulans as a reference species for cell biology.

Theobald, Sebastian [Technical University of Denma↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii

Tilletia maclaganii is a smut fungal pathogen that causes significant biomass reduction of switchgrass ( Panicum virgatum ) used for animal forage and biofuel production. Here we present the annotated genome of T. maclaganii , strain Tm001-NY21, estimated at 42.79 Mb in size, in 53 assembled contigs and encoding 10,235 predicted genes. This genome will be important for future comparative studies of Ustilaginales across its geographic and host range.

PacBio↗

The architecture of resilience: a genome assembly of Myrothamnus flabellifolia sheds light on desiccation tolerance and sex determination

Myrothamnus flabellifolia is a dioecious resurrection plant endemic to southern Africa that has become an important model for understanding desiccation tolerance. Despite its ecological and medicinal significance, genomic and transcriptomic resources for the species are limited. We generated a chromosome-level, haplotype-resolved reference genome assembly and annotation for M. flabellifolia and conducted transcriptomic profiling across a natural dehydration–rehydration time course in the field. Genome architecture and sex determination were characterized, and co-expression network and cis-regulatory element (CRE) enrichment analyses were used to investigate dynamic responses to desiccation. The 1.28-Gb genome exhibits unusually consistent chromatin architecture with unique chromosome organization across highly divergent haplotypes. We identified an XY sexual system with a small sex-determining region on Chromosome 8. Transcriptomic responses varied with dehydration severity, pointing to early suppression of growth, progressive activation of protective mechanisms, and subsequent return to homeostasis upon rehydration. Late embryogenesis abundant and early light-induced protein transcripts were dynamically regulated and showed enrichment of abscisic acid and stress-responsive CREs pointing toward conserved responses. Together, this study provides foundational resources for understanding the genomic architecture and reproductive biology of M. flabellifolia and offers new insights into the mechanisms of desiccation tolerance.

chromosome structure↗

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag↗

Multi‐season analysis reveals hundreds of drought‐responsive genes in sorghum

Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.

Cole, Benjamin [USDOE Joint Genome Institute (JGI)↗

Transcriptomic Analysis of the CAM Species Kalanchoë fedtschenkoi Under Low- and High-Temperature Regimes

Temperature stress is one of the major limiting environmental factors that negatively impact global crop yields. Kalanchoë fedtschenkoi is an obligate crassulacean acid metabolism (CAM) plant species, exhibiting much higher water-use efficiency and tolerance to drought and heat stresses than C 3 or C 4 plant species. Previous studies on gene expression responses to low- or high-temperature stress have been focused on C 3 and C 4 plants. There is a lack of information about the regulation of gene expression by low and high temperatures in CAM plants. To address this knowledge gap, we performed transcriptome sequencing (RNA-Seq) of leaf and root tissues of K. fedtschenkoi under cold (8 °C), normal (25 °C), and heat (37 °C) conditions at dawn (i.e., 2 h before the light period) and dusk (i.e., 2 h before the dark period). Our analysis revealed differentially expressed genes (DEGs) under cold or heat treatment in comparison to normal conditions in leaf or root tissue at each of the two time points. In particular, DEGs exhibiting either the same or opposite direction of expression change (either up-regulated or down-regulated) under cold and heat treatments were identified. In addition, we analyzed gene co-expression modules regulated by cold or heat treatment, and we performed in-depth analyses of expression regulation by temperature stresses for selected gene categories, including CAM-related genes, genes encoding heat shock factors and heat shock proteins, circadian rhythm genes, and stomatal movement genes. Our study highlights both the common and distinct molecular strategies employed by CAM and C 3 /C 4 plants in adapting to extreme temperatures, providing new insights into the molecular mechanisms underlying temperature stress responses in CAM species.

59 BASIC BIOLOGICAL SCIENCES↗

Structure and sequence evolution in the pennycress ( Thlaspi arvense ) pangenome

Eukaryotic genomes harbor many forms of variation, including nucleotide diversity and structural polymorphisms, which experience natural selection and contribute to genome evolution and biodiversity. Harnessing this variation for agriculture hinges on our ability to detect, quantify, catalog, and deploy genetic diversity. Here, we explore seven complete genomes of the emerging biofuel crop pennycress ( Thlaspi arvense ) drawn from across the species' current genetic diversity to catalog variation in genome structure and content. Across this new pangenome resource, we find contrasting evolutionary modes in different genomic zones. Gene-poor, repeat-rich pericentromeric regions experience frequent rearrangements, including repeated centromere repositioning. By contrast, conserved gene-dense chromosome arms maintain large-scale synteny across accessions even in fast-evolving NOD-like receptor immune genes, where microsynteny breaks down across species, but gene cluster positioning macrosynteny is maintained. Our findings highlight that multiple elements of the genome experience dynamic evolution that conserves functional content on the chromosome scale but allows repositioning and presence–absence variation on a local scale. This diversity is invisible to classical reference-based strategies and highlights the strength and utility of pangenomic resources. These results provide a valuable case study of rapid genomic structural evolution within a species and powerful resources for crop development in an emerging biofuel crop.

Thlaspi arvense↗

Rhythmic Mechanisms Governing CAM Photosynthesis in Kalanchoe fedtschenkoi : High-Resolution Temporal Transcriptomics

Crassulacean acid metabolism (CAM) is a specialized photosynthetic pathway that enhances water-use efficiency by temporally separating nocturnal CO 2 uptake from daytime decarboxylation and carbon fixation. To uncover the regulatory mechanisms coordinating these temporal dynamics, we generated high-resolution, 48 h time-course transcriptomes for the CAM model Kalanchoe fedtschenkoi under both 12 h/12 h light/dark (LD) cycles and continuous light (LL). A rhythmicity analysis revealed that diel light cues are the dominant driver of transcript oscillations: 16,810 genes (54.3% of annotated genes) exhibited rhythmic expression only under LD, whereas just 399 genes (1.3%) remained rhythmic under LL. A smaller set of 3009 genes (9.7%) oscillated in both conditions, indicating that the intrinsic circadian clock sustains rhythmicity for a limited subset of the transcriptome. A gene co-expression network analysis revealed extensive integration between circadian clock components, core CAM pathway enzymes, and stomatal regulators, defining regulatory modules that coordinate metabolic and physiological timing. Notably, key hub genes associated with post-translational and post-transcriptional regulation, including the E3 ubiquitin ligase HUB2 and several pentatricopeptide repeat (PPR) proteins, act as central nodes in CAM-associated networks. This discovery implicates epigenetic and organellar regulation as previously unrecognized critical tiers of control in CAM. Together, our results support a regulatory model in which CAM rhythmicity is governed by both external light/dark cues and the endogenous circadian clock through multi-level control spanning transcriptional and protein-level regulation. To support community exploration, we also provide an interactive eFP (electronic Fluorescent Pictograph) browser for visualizing time-resolved gene expression profiles.

09 BIOMASS FUELS↗

Three pairs of fungal Trametes strains isolated from distinct geographic origins show conserved genomic features and adaptive response to plant biomass

The genomes of white-rot fungi hold extended repertoires of enzymes active on virtually all the chemical bonds that intertwine lignocellulose polymers, and several Trametes species have been identified as powerful tools for biorefinery or bioremediation. However, only few studies have addressed the intra-species polymorphism one would expect from fungal strains collected in contrasted environments. We compared the genome sequence of pairs of strains collected in different geographic areas, for each of three fungal species. Using an updated list of the predicted functions for fungal ligno- and cellulolytic enzymes (CAZymes), we observed a high conservation of the gene repertoires among the six strains. We compared the adaptative response of the fungi grown on crystalline cellulose, wheat straw, aspen or pine sawdust by transcriptomics and secretomics. The gene regulation profiles were determined by the species and the substrates, rather than the strain. The secretomes did not show marked differences in the sets of secreted CAZymes after 3 day-growth on the substrates. We identified five transcription factor genes and two sesquiterpenoid synthesis genes induced during growth on lignocellulose. Wider studies using larger sets of strains will be necessary to evaluate the genericity of our findings, and to assess the phenotype diversity one could expect from geographic diversity as compared to taxonomic diversity in Trametes fungi.

Drula, E. [French National Research Institute for ↗

Label-free structural imaging of plant roots and microbes using third-harmonic generation microscopy

Root biology is pivotal in addressing global challenges including sustainable agriculture and climate change. However, roots have been relatively understudied among plant organs, partly due to the difficulties in imaging root structures in their natural environment. Here we used microfabricated ecosystems (EcoFABs) to establish growing environments with optical access and employed nonlinear multimodal microscopy of third-harmonic generation (THG) and three-photon fluorescence (3PF) to achieve label-free, in situ imaging of live roots and microbes at high spatiotemporal resolution. THG enabled us to observe key plant root structures including the vasculature, Casparian strips, dividing meristematic cells, and root cap cells, as well as subcellular features including nuclear envelopes, nucleoli, starch granules, and putative stress granules. THG from the cell walls of bacteria and fungi also provides label-free contrast for visualizing these microbes in the root rhizosphere. With simultaneously recorded 3PF signal, we demonstrated our ability to investigate root-microbe interactions by achieving single-bacterium tracking and subcellular imaging of fungal spores and hyphae in the rhizosphere.

Pan, Daisong [University of California, Berkeley, ↗

ENVnet provides a global molecular resource of dissolved organic matter

Dissolved organic matter (DOM) is an important component of Earth's carbon cycle and one of the planet's most chemically diverse pools, yet the molecular structures of its constituents remain largely unresolved. This limitation has hindered our ability to link DOM composition to microbial processes and ecosystem function. Here we present ENVnet, a global molecular repository built from tandem mass spectrometry data collected across 13 terrestrial and aquatic environment types, including 419 newly generated samples that expand publicly available DOM metabolomics data and cover previously underrepresented environments. By computationally deconvolving chimeric mass spectra, a longstanding challenge in environmental metabolomics, we recover high-quality fragmentation data for >22,000 distinct molecular features (defined by a specific precursor mass and fragmentation pattern). Using ENVnet, we uncover conserved and environment-specific molecular patterns in DOM composition and underlying biogeochemical processes. We also use molecular features encoded in ENVnet to train predictive models of DOM persistence, allowing molecular-level assessment of microbial turnover in independent systems.

54 ENVIRONMENTAL SCIENCES↗