Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Genomic prediction of regional-scale performance in switchgrass ( Panicum virgatum ) by accounting for genotype-by-environment variation and yield surrogate traits

Switchgrass is a potential crop for bioenergy or carbon capture schemes, but further yield improvements through selective breeding are needed to encourage commercialization. To identify promising switchgrass germplasm for future breeding efforts, we conducted multisite and multitrait genomic prediction with a diversity panel of 630 genotypes from 4 switchgrass subpopulations (Gulf, Midwest, Coastal, and Texas), which were measured for spaced plant biomass yield across 10 sites. Our study focused on the use of genomic prediction to share information among traits and environments. Specifically, we evaluated the predictive ability of cross-validation (CV) schemes using only genetic data and the training set (cross-validation 1: CV1), a subset of the sites (cross-validation 2: CV2), and/or with 2 yield surrogates (flowering time and fall plant height). We found that genotype-by-environment interactions were largely due to the north–south distribution of sites. The genetic correlations between the yield surrogates and the biomass yield were generally positive (mean height r = 0.85; mean flowering time r = 0.45) and did not vary due to subpopulation or growing region (North, Middle, or South). Genomic prediction models had CV predictive abilities of –0.02 for individuals using only genetic data (CV1), but 0.55, 0.69, 0.76, 0.81, and 0.84 for individuals with biomass performance data from 1, 2, 3, 4, and 5 sites included in the training data (CV2), respectively. To simulate a resource-limited breeding program, we determined the predictive ability of models provided with the following: 1 site observation of flowering time (0.39); 1 site observation of flowering time and fall height (0.51); 1 site observation of fall height (0.52); 1 site observation of biomass (0.55); and 5 site observations of biomass yield (0.84). The ability to share information at a regional scale is very encouraging, but further research is required to accurately translate spaced plant biomass to commercial-scale sward biomass performance.

09 BIOMASS FUELS↗

Genome evolution and transcriptome plasticity is associated with adaptation to monocot and dicot plants in Colletotrichum fungi

Colletotrichum fungi infect a wide diversity of monocot and dicot hosts, causing diseases on almost all economically important plants worldwide. Colletotrichum is also a suitable model for studying gene family evolution on a fine scale to uncover events in the genome associated with biological changes. Here we present the genome sequences of 30 Colletotrichum species covering the diversity within the genus. Evolutionary analyses revealed that the Colletotrichum ancestor diverged in the late Cretaceous in parallel with the diversification of flowering plants. We provide evidence of independent host jumps from dicots to monocots during the evolution of Colletotrichum, coinciding with a progressive shrinking of the plant cell wall degradative arsenal and expansions in lineage-specific gene families. Comparative transcriptomics of 4 species adapted to different hosts revealed similarity in gene content but high diversity in the modulation of their transcription profiles on different plant substrates. Combining genomics and transcriptomics, we identified a set of core genes such as specific transcription factors, putatively involved in plant cell wall degradation. These results indicate that the ancestral Colletotrichum were associated with dicot plants and certain branches progressively adapted to different monocot hosts, reshaping the gene content and its regulation.

59 BASIC BIOLOGICAL SCIENCES↗

Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing

Abstract The microbiome is a complex community of microorganisms, encompassing prokaryotic (bacterial and archaeal), eukaryotic, and viral entities. This microbial ensemble plays a pivotal role in influencing the health and productivity of diverse ecosystems while shaping the web of life. However, many software suites developed to study microbiomes analyze only the prokaryotic community and provide limited to no support for viruses and microeukaryotes. Previously, we introduced the Viral Eukaryotic Bacterial Archaeal (VEBA) open-source software suite to address this critical gap in microbiome research by extending genome-resolved analysis beyond prokaryotes to encompass the understudied realms of eukaryotes and viruses. Here we present VEBA 2.0 with key updates including a comprehensive clustered microeukaryotic protein database, rapid genome/protein-level clustering, bioprospecting, non-coding/organelle gene modeling, genome-resolved taxonomic/pathway profiling, long-read support, and containerization. We demonstrate VEBA’s versatile application through the analysis of diverse case studies including marine water, Siberian permafrost, and white-tailed deer lung tissues with the latter showcasing how to identify integrated viruses. VEBA represents a crucial advancement in microbiome research, offering a powerful and accessible software suite that bridges the gap between genomics and biotechnological solutions.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic analysis of Klebsiella aerogenes circulating in New Mexico

Klebsiella aerogenes is an opportunistic pathogen and a growing cause of healthcare-associated infections, characterized by multidrug resistance and the emergence of global high-risk clones. However, regional genomic surveillance data remain limited. Here, we sought to characterize the population structure, transmission dynamics and resistance mechanisms of clinical K. aerogenes in Albuquerque, New Mexico. We sequenced 177 clinical isolates collected between 2021 and 2023. We also developed a novel, species-specific PopPUNK database to facilitate rapid, high-resolution typing. The New Mexico K. aerogenes population was diverse but dominated by two global pandemic lineages, ST93 (47.5%) and ST4 (7.9%), which were significantly enriched for the virulence factors yersiniabactin and colibactin. Genomic evidence for recent local transmission was rare, with only four putative transmission pairs identified. The resistome was characterized by intrinsic and adaptive mutations. Nearly all isolates possessed gyrA mutations associated with decreased fluoroquinolone susceptibility. Mutations in the AmpC regulator AmpD and the outer membrane porin Omp36 were common, particularly within the dominant ST93 lineage. These mutations have been associated with increased AmpC-mediated carbapenem resistance. Our findings underscore the critical importance of genomic surveillance to monitor the transmission and evolution of adaptive resistance.

59 BASIC BIOLOGICAL SCIENCES↗

High-quality draft genome sequence of Thermobifida halotolerans DSM 44931

Here, we report the genome sequence of Thermobifida halotolerans DSM 44931, a bacterium that was originally isolated from a salt mine in the Yunnan Province of China. This genome was sequenced using Pacific Biosciences sequencing technology and was assembled into 2 contigs in 2 scaffolds. It has a total length of 5,506,851 bp and a GC content of 71.16%. Functional annotation of this genome provides further metabolic insight into this species.

actinomycete↗

Metagenome-assembled genomes measured at 3 depths during snowmelt period in East River, CO (March, May, and June, September 2017)

Snowmelt is a critical biogeochemical period that accounts for large nitrogen (N) export events from high-elevation watersheds. Soil microbial populations bloom and immobilize N during snowmelt, yet the population size crashes in spring, which releases a pulse of soil N. We sought to discover the N sources fueling this microbial bloom and determine the fate of N following microbial die-off. Here, focusing on the snowmelt period within a headwater catchment of the Upper Colorado River Basin (East River, CO), we deployed strain-resolved metagenomics to identify the metabolic pathways and processes that mobilize soil N during and after snowmelt. Soil metagenome samples were taken from 6 snowpits from 3 depths (0-5cm, 5-15cm, >15cm) at 4 time points during snowmelt period (March 2017, May 2017, and June 2017, September 2017) generating 48 metagenomes. We reconstructed 474 metagenome-assembled genomes (MAGs) across all metagenomes.All 48 metagenomes were sequenced at JGI and raw data can be found under JGI (Joint Genome Institute) GOLD Study Gs0135149. Metagenome assemblies from IMG under the same study were used for genome binning. This dataset (1) a zip file of 474 MAGs (as fasta files, Gs0135149_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0135149.kml), (4) metagenome metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (metagenomes.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Fermented Foods Microbial Genomes Database

This database contains ~4,300 microbial genomes assembled from diverse fermented foods. These genomes were obtained from a larger set of 13,850 microbial genomes by clustering them at 99% average nucleotide identity (ANI) to create a "species"-representative database.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic and Phenotypic Characterization of Yeast Biosensor for Deep-space Radiation

The BioSentinel mission was selected to launch as a secondary payload onboard NASA Exploration Mission 1 (EM-1) in 2018. In BioSentinel, the budding yeast Saccharomyces cerevisiae will be used as a biosensor to measure the long-term impact of deep-space radiation to living organisms. In the 4U-payload, desiccated yeast cells from different strains will be stored inside microfluidic cards equipped with 3-color LED optical detection system to monitor cell growth and metabolic activity. At different times throughout the 12-month mission, these cards will be filled with liquid yeast growth media to rehydrate and grow the desiccated cells. The growth and metabolic rates of wild-type and radiation-sensitive strains in deep-space radiation environment will be compared to the rates measured in the ground- and microgravity-control units. These rates will also be correlated with measurements obtained from onboard physical dosimeters. In our preliminary long-term desiccation study, we found that air-drying yeast cells in 10% trehalose is the best method of cell preservation in order to survive the entire 18-month mission duration (6-month pre-launch plus 12-month full-mission periods). However, our study also revealed that desiccated yeast cells have decreasing viability over time when stored in payload-like environment. This suggests that the yeast biosensor will have different population of cells at different time points during the long-term mission. In this study, we are characterizing genomic and phenotypic changes in our yeast biosensor due to long-term storage and desiccation. For each yeast strain that will be part of the biosensor, several clones were reisolated after long-term storage by desiccation. These clones were compared to their respective original isolate in terms of genomic composition, desiccation tolerance and radiation sensitivity. Interestingly, clones from a radiation-sensitive mutant have better desiccation tolerance compared to their original isolate without losing radiation sensitivity. We employed Next-Generation Sequencing technology to better understand this phenotypic variation. Current effort is focusing on the analysis of high-throughput sequencing data to look for genomic changes in these reisolated clones compared to their original isolate.

yeast↗

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗

Sequential membrane- and protein-bound organelles compartmentalize genomes during phage infection

Many eukaryotic viruses require membrane-bound compartments for replication, but no such organelles are known to be formed by prokaryotic viruses. Bacteriophages of the Chimalliviridae family sequester their genomes within a phage-generated organelle, the phage nucleus, which is enclosed by a lattice of the viral protein ChmA. We show that inhibiting phage nucleus formation arrests infections at an early stage in which the injected phage genome is enclosed within a membrane-bound early phage infection (EPI) vesicle. Early phage genes are expressed from the EPI vesicle, demonstrating its functionality as a prokaryotic, transcriptionally active, membrane-bound organelle. We also show that the phage nucleus is essential, with genome replication beginning after the injected DNA is transferred from the EPI vesicle to the phage nucleus. Our results show that Chimalliviridae require two sophisticated subcellular compartments of distinct compositions and functions that facilitate successive stages of the viral life cycle.

59 BASIC BIOLOGICAL SCIENCES↗

CRISPRi-ART enables functional genomics of diverse bacteriophages using RNA-binding dCas13d

Bacteriophages constitute one of the largest reservoirs of genes of unknown function in the biosphere. Even in well-characterized phages, the functions of most genes remain unknown. Experimental approaches to study phage gene fitness and function at genome scale are lacking, partly because phages subvert many modern functional genomics tools. Here we leverage RNA-targeting dCas13d to selectively interfere with protein translation and to measure phage gene fitness at a transcriptome-wide scale. We find CRISPR Interference through Antisense RNA-Targeting (CRISPRi-ART) to be effective across phage phylogeny, from model ssRNA, ssDNA and dsDNA phages to nucleus-forming jumbo phages. Using CRISPRi-ART, we determine a conserved role of diverse rII homologues in subverting phage Lambda RexAB-mediated immunity to superinfection and identify genes critical for phage fitness. CRISPRi-ART establishes a broad-spectrum phage functional genomics platform, revealing more than 90 previously unknown genes important for phage fitness.

59 BASIC BIOLOGICAL SCIENCES↗

Targeted delivery of genome editors in vivo

Genome editing has revolutionized the treatment of genetic diseases, yet the difficulty of tissue-specific delivery currently limits applications of editing technology. Here, in this Review, we discuss preclinical and clinical advances in delivering genome editors with both established and emerging delivery mechanisms. Targeted delivery promises to considerably expand the therapeutic applicability of genome editing, moving closer to the ideal of a precise ‘magic bullet’ that safely and effectively treats diverse genetic disorders.

Ngo, Wayne [University of California, Berkeley, CA↗

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Testing for the Genomic Footprint of Conflict Between Life Stages in an Angiosperm and Moss Species

Abstract The maintenance of genetic variation by balancing selection is of considerable interest to evolutionary biologists. An important but understudied potential driver of balancing selection is antagonistic pleiotropy between diploid and haploid stages of the plant life cycle. Despite sharing a common genome, sporophytes (2n) and gametophytes (n) may undergo differential or even opposing selection. Theoretical work suggests antagonistic pleiotropy between life stages can generate balancing selection and maintain genetic variation. Despite the potential for far-reaching consequences of gametophytic selection, empirical tests of its pleiotropic effects (neutral, synergistic, or antagonistic) on sporophytes are generally lacking. Here, we examined the population genomic signals of selection across life stages in the angiosperm Rumex hastatulus and the moss Ceratodon purpureus. We compared gene expression between life stages and sexes, combined with neutral diversity statistics and the analysis of the distribution of fitness effects. In contrast to what would be predicted under balancing selection due to antagonistic pleiotropy, we found that unbiased genes between life stages were under stronger purifying selection, likely explained by a predominance of synergistic pleiotropy between life stages and strong purifying selection on broadly expressed genes. In addition, we found that 30% of candidate genes under balancing selection in R. hastatulus were located within inversion polymorphisms. Our findings provide novel insights into the genome-wide characteristics and consequences of plant gametophytic selection.

Evolutionary Biology↗

Structural basis for a highly conserved RNA-mediated enteroviral genome replication

Abstract Enteroviruses contain conserved RNA structures at the extreme 5′ end of their genomes that recruit essential proteins 3CD and PCBP2 to promote genome replication. However, the high-resolution structures and mechanisms of these replication-linked RNAs (REPLRs) are limited. Here, we determined the crystal structures of the coxsackievirus B3 and rhinoviruses B14 and C15 REPLRs at 1.54, 2.2 and 2.54 Å resolution, revealing a highly conserved H-type four-way junction fold with co-axially stacked sA-sD and sB-sC helices that are stabilized by a long-range A•C•U base-triple. Such conserved features observed in the crystal structures also allowed us to predict the models of several other enteroviral REPLRs using homology modeling, which generated models almost identical to the experimentally determined structures. Moreover, our structure-guided binding studies with recombinantly purified full-length human PCBP2 showed that two previously proposed binding sites, the sB-loop and 3′ spacer, reside proximally and bind a single PCBP2. Additionally, the DNA oligos complementary to the 3′ spacer, the high-affinity PCBP2 binding site, abrogated its interactions with enteroviral REPLRs, suggesting the critical roles of this single-stranded region in recruiting PCBP2 for enteroviral genome replication and illuminating the promising prospects of developing therapeutics against enteroviral infections targeting this replication platform.

Biochemistry & Molecular Biology↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27↗