Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Gene family expansions and transcriptome signatures uncover fungal adaptations to wood decay

Because they comprise some of the most efficient wood-decayers, Polyporales fungi impact carbon cycling in forest environment. Despite continuous discoveries on the enzymatic machinery involved in wood decomposition, the vision on their evolutionary adaptation to wood decay and genome diversity remains incomplete. We combined the genome sequence information from 50 Polyporales species, including 26 newly sequenced genomes and sought for genomic and functional adaptations to wood decay through the analysis of genome composition and transcriptome responses to different carbon sources. The genomes of Polyporales from different phylogenetic clades showed poor conservation in macrosynteny, indicative of genome rearrangements. We observed different gene family expansion/contraction histories for plant cell wall degrading enzymes in core polyporoids and phlebioids and captured expansions for genes involved in signalling and regulation in the lineages of white rotters. Furthermore, we identified conserved cupredoxins, thaumatin-like proteins and lytic polysaccharide monooxygenases with a yet uncharacterized appended module as new candidate players in wood decomposition. Given the current need for enzymatic toolkits dedicated to the transformation of renewable carbon sources, the observed genomic diversity among Polyporales strengthens the relevance of mining Polyporales biodiversity to understand the molecular mechanisms of wood decay.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic Determinants of Microbial Survival in Space

Space flight agencies envision a future for humankind beyond Earth, including missions back to the Moon and to Mars in the coming decades. Sending humans into space inevitably includes their microbiomes as well, leading to trillions of bacteria being shed in their living areas. These bacteria shape the lives of their hosts as well as their environment; thus, it is crucial to understand the adaptations of these microbial spacefarers in spaceflight conditions. We aimed to elucidate the genetic determinants of microbial survival in space using a pan-genome analysis of 12 genera cultured from the International Space Station (ISS) from 2017 to 2018. Analysis was performed on each of the genera individually with terrestrial analogs to identify the core and accessory genomes of the spaceflight and terrestrial strains. We then compared the flight and terrestrial core and accessory genomes for each genera using a Bray-Curtis index and visualized the resulting dissimilarity using an Non-Metric Dimensional Scaling plot. The core proteins available in only the spaceflight organisms were then manually characterized for function and genomic location. In every core genome comparison in each genus, there was significant dissimilarity in the core of the spaceflight organisms when compared to the terrestrial organisms. This trend was present in some of the accessory genomes, but was not ubiquitous. Functional analysis of the core content of the ISS genomes showed the majority of genes unique to the core were clustered by location. These gene clusters suggested a set of genetic determinants confer survival in spacecraft-built environments, notably through the uptake of extracellular DNA such as bacteriophage and plasmids. The clear difference between spaceflight and terrestrial microorganisms shows that spaceflight conditions are selective, which has long term implications for their human hosts and environments.

MoBE↗

Leveraging computational genomics to understand the molecular basis of metal homeostasis

Genome-based data is helping to reveal the diverse strategies plants and algae use to maintain metal homeostasis. In addition to acquisition, distribution and storage of metals, acclimating to feast or famine can involve a wealth of genes that we are just now starting to understand. The fast-paced acquisition of genome-based data, however, is far outpacing our ability to experimentally characterize protein function. Computational genomic approaches are needed to fill the gap between what is known and unknown. To avoid misconstruing bioinformatically derived data, which is the root cause of the inaccurate functional annotations that plague databases, functional inferences from diverse sources and contextualization of that evidence with a robust understanding of protein family evolution is needed. Phylogenomic- and comparative-genomic-based studies can aid in the interpretation of experimental data or provide a spark for the discovery of a new function. These analyses not only lead to novel insight into a target protein's function but can generate thought-provoking insights across protein families.

59 BASIC BIOLOGICAL SCIENCES↗

Four chromosome scale genomes and a pan-genome annotation to accelerate pecan tree breeding

Genome-enabled biotechnologies have the potential to accelerate breeding efforts in long-lived perennial crop species. Despite the transformative potential of molecular tools in pecan and other outcrossing tree species, highly heterozygous genomes, significant presence–absence gene content variation, and histories of interspecific hybridization have constrained breeding efforts. To overcome these challenges, here, we present diploid genome assemblies and annotations of four outbred pecan genotypes, including a PacBio HiFi chromosome-scale assembly of both haplotypes of the ‘Pawnee’ cultivar. Comparative analysis and pan-genome integration reveal substantial and likely adaptive interspecific genomic introgressions, including an over-retained haplotype introgressed from bitternut hickory into pecan breeding pedigrees. Further, by leveraging our pan-genome presence–absence and functional annotation database among genomes and within the two outbred haplotypes of the ‘Lakota’ genome, we identify candidate genes for pest and pathogen resistance. Combined, these analyses and resources highlight significant progress towards functional and quantitative genomics in highly diverse and outbred crops.

54 ENVIRONMENTAL SCIENCES↗

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

Tunturi virus isolates and metagenome-assembled viral genomes provide insights into the virome of Acidobacteriota in Arctic tundra soils

Arctic soils are climate-critical areas, where microorganisms play crucial roles in nutrient cycling processes. Acidobacteriota are phylogenetically and physiologically diverse bacteria that are abundant and active in Arctic tundra soils. Still, surprisingly little is known about acidobacterial viruses in general and those residing in the Arctic in particular. Here, we applied both culture-dependent and -independent methods to study the virome of Acidobacteriota in Arctic soils. Five virus isolates, Tunturi 1–5, were obtained from Arctic tundra soils, Kilpisjärvi, Finland (69°N), using Tunturiibacter spp. strains originating from the same area as hosts. The new virus isolates have tailed particles with podo- (Tunturi 1, 2, 3), sipho- (Tunturi 4), or myovirus-like (Tunturi 5) morphologies. The dsDNA genomes of the viral isolates are 63–98 kbp long, except Tunturi 5, which is a jumbo phage with a 309-kbp genome. Tunturi 1 and Tunturi 2 share 88% overall nucleotide identity, while the other three are not related to one another. For over half of the open reading frames in Tunturi genomes, no functions could be predicted. To further assess the Acidobacteriota-associated viral diversity in Kilpisjärvi soils, bulk metagenomes from the same soils were explored and a total of 1881 viral operational taxonomic units (vOTUs) were bioinformatically predicted. Almost all vOTUs (98%) were assigned to the class Caudoviricetes. For 125 vOTUs, including five (near-)complete ones, Acidobacteriota hosts were predicted. Acidobacteriota-linked vOTUs were abundant across sites, especially in fens. Terriglobia-associated proviruses were observed in Kilpisjärvi soils, being related to proviruses from distant soils and other biomes. Approximately genus- or higher-level similarities were found between the Tunturi viruses, Kilpisjärvi vOTUs, and other soil vOTUs, suggesting some shared groups of Acidobacteriota viruses across soils. This study provides acidobacterial virus isolates as laboratory models for future research and adds insights into the diversity of viral communities associated with Acidobacteriota in tundra soils. Predicted virus-host links and viral gene functions suggest various interactions between viruses and their host microorganisms. Largely unknown sequences in the isolates and metagenome-assembled viral genomes highlight a need for more extensive sampling of Arctic soils to better understand viral functions and contributions to ecosystem-wide cycling processes in the Arctic.

54 ENVIRONMENTAL SCIENCES↗

IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata

Viruses are widely recognized as critical members of all microbiomes. Metagenomics enables large-scale exploration of the global virosphere, progressively revealing the extensive genomic diversity of viruses on Earth and highlighting the myriad of ways by which viruses impact biological processes. IMG/VR provides access to the largest collection of viral sequences obtained from (meta)genomes, along with functional annotation and rich metadata. A web interface enables users to efficiently browse and search viruses based on genome features and/or sequence similarity. Here, for this work, we present the fourth version of IMG/VR, composed of >15 million virus genomes and genome fragments, a ≈6-fold increase in size compared to the previous version. These clustered into 8.7 million viral operational taxonomic units, including 231 408 with at least one high-quality representative. Viral sequences in IMG/VR are now systematically identified from genomes, metagenomes, and metatranscriptomes using a new detection approach (geNomad), and IMG standard annotation are complemented with genome quality estimation using CheckV, taxonomic classification reflecting the latest taxonomic standards, and microbial host taxonomy prediction. IMG/VR v4 is available at https://img.jgi.doe.gov/vr, and the underlying data are available to download at https://genome.jgi.doe.gov/portal/IMG_VR.

59 BASIC BIOLOGICAL SCIENCES↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Stable hypermutators revealed by the genomic landscape of genes involved in genome stability among yeast species

Mutator phenotypes are short-lived due to the rapid accumulation of deleterious mutations. Yet, recent observations reveal that certain fungi can undergo prolonged accelerated evolution after losing genes involved in DNA repair. Here, we surveyed 1,154 yeast genomes representing nearly all known yeast species of the subphylum Saccharomycotina (phylum Ascomycota) to examine the relationship between reduced gene repertoires broadly associated with genome stability functions (e.g., DNA repair, cell cycle) and elevated evolutionary rates. We identified three distantly related lineages—encompassing 12% of species—that had both the most streamlined sets of genes involved in genome stability (specifically DNA repair) and the highest evolutionary rates in the entire subphylum. Two of these “faster-evolving lineages” (FELs)—a subclade within the order Pichiales and the Wickerhamiella/Starmerella (W/S) clade (order Dipodascales)—are described here for the first time, while the third corresponds to a previously documented Hanseniaspora FEL. Examination of genome stability gene repertoires revealed a set of genes predominantly absent in these three FELs, suggesting a potential role in the observed acceleration of evolutionary rates. In the W/S clade, genomic signatures are consistent with a substantial mutational burden, including pronounced A|T bias and endogenous DNA damage. Interestingly, we found that the W/S clade also contains DNA repair genes possibly acquired through horizontal gene transfer, including a photolyase of bacterial origin. These findings highlight how hypermutators can persist across macroevolutionary timescales, potentially linked to the loss of genes related with genome stability, with horizontal gene transfer as a possible avenue for partial functional compensation.

DNA repair↗

Modeling the Influenza A NP-vRNA-Polymerase Complex in Atomic Detail

Seasonal flu is an acute respiratory disease that exacts a massive toll on human populations, healthcare systems and economies. The disease is caused by an enveloped Influenza virus containing eight ribonucleoprotein (RNP) complexes. Each RNP incorporates multiple copies of nucleoprotein (NP), a fragment of the viral genome (vRNA), and a viral RNA-dependent RNA polymerase (POL), and is responsible for packaging the viral genome and performing critical functions including replication and transcription. A complete model of an Influenza RNP in atomic detail can elucidate the structural basis for viral genome functions, and identify potential targets for viral therapeutics. In this work we construct a model of a complete Influenza A RNP complex in atomic detail using multiple sources of structural and sequence information and a series of homology-modeling techniques, including a motif-matching fragment assembly method. Our final model provides a rationale for experimentally-observed changes to viral polymerase activity in numerous mutational assays. Further, our model reveals specific interactions between the three primary structural components of the RNP, including potential targets for blocking POL-binding to the NP-vRNA complex. The methods developed in this work open the possibility of elucidating other functionally-relevant atomic-scale interactions in additional RNP structures and other biomolecular complexes.

59 BASIC BIOLOGICAL SCIENCES↗

MultiPhATE2: code for functional annotation and comparison of phage genomes

To address a need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of multiPhATE, a functional annotation code released previously, multiPhATE2 performs gene finding using multiple algorithms, compares the results of the algorithms, performs functional annotation of coding sequences, and incorporates additional search algorithms and databases to extend the search space of the original code. MultiPhATE2 performs gene matching among sets of closely related bacteriophage genomes, and uses multiprocessing to speed computations. MultiPhATE2 can be re-started at multiple points within the workflow to allow the user to examine intermediate results and adjust the subsequent computations accordingly. In addition, multiPhATE2 accommodates custom gene calls and sequence databases, again adding flexibility. MultiPhATE2 was implemented in Python 3.7 and runs as a command-line code under Linux or MAC operating systems. Full documentation is provided as a README file and a Wiki website.

59 BASIC BIOLOGICAL SCIENCES↗

Genomics-enabled analysis of specialized metabolism in bioenergy crops: Current progress and challenges

Plants produce a staggering diversity of specialized small molecule metabolites that play vital roles in mediating environmental interactions and stress adaptation. This chemical diversity derives from dynamic biosynthetic pathway networks that are often species-specific and operate under tight spatiotemporal and environmental control. A growing divide between demand and environmental challenges in food and bioenergy crop production have intensified research on these complex metabolite networks and their contribution to crop fitness. High-throughput omics technologies provide access to ever-increasing data resources for investigating plant metabolism. However, the efficiency of using such system-wide data to decode the gene and enzyme functions controlling specialized metabolism has remained limited; due largely to the recalcitrance of many plants to genetic approaches and the lack of ‘user-friendly’ biochemical tools for studying the diverse enzyme classes involved in specialized metabolism. With emphasis on terpenoid metabolism in the bioenergy crop switchgrass as an example, this review aims to illustrate current advances and challenges in the application of DNA synthesis and synthetic biology tools for accelerating the functional discovery of genes, enzymes and pathways in plant specialized metabolism. These technologies have accelerated knowledge development on the biosynthesis and physiological roles of diverse metabolite networks across many ecologically and economically important plant species and can provide resources for application to precision breeding and natural product metabolic engineering.

59 BASIC BIOLOGICAL SCIENCES↗

High Throughput Genome Releaser

In this study, we present the development of a High Throughput Genome Releaser, an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed for rapid, cost-effective, and efficient DNA extraction, optimized for subsequent PCR reactions. Our experimentation with various synthetic materials led us to select a particular type of plastic that mirrors the properties of glass cover slides, providing a smooth surface and effective compression capabilities. We engineered a 96-well device equipped with a 96-well plate and a top rod, operable both manually and automatically, which is compatible with widely used liquid-handling robot decks. This compatibility enhances ease of use in high-throughput PCR setups. Additionally, we developed software to support its automatic functions. The genome releaser facilitates the extraction of PCR-amplifiable genomic DNA from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells. This versatility could significantly advance biomanufacturing processes.

42 ENGINEERING↗

Systematic discovery of pseudomonad genetic factors involved in sensitivity to tailocins

Tailocins are bactericidal protein complexes produced by a wide variety of bacteria that kill closely related strains and may play a role in microbial community structure. Thanks to their high specificity, tailocins have been proposed as precision antibacterial agents for therapeutic applications. Compared to tailed phages, with whom they share an evolutionary and morphological relationship, bacterially produced tailocins kill their host upon production but producing strains display resistance to self-intoxication. Though lipopolysaccharide (LPS) has been shown to act as a receptor for tailocins, the breadth of factors involved in tailocin sensitivity, and the mechanisms behind resistance to self-intoxication, remain unclear. Here, we employed genome-wide screens in four non-model pseudomonads to identify mutants with altered fitness in the presence of tailocins produced by closely related pseudomonads. Our mutant screens identified O-antigen composition and display as most important in defining sensitivity to our tailocins. In addition, the screens suggest LPS thinning as a mechanism by which resistant strains can become more sensitive to tailocins. Furthermore, we validate many of these novel findings, and extend these observations of tailocin sensitivity to 130 genome-sequenced pseudomonads. This work offers insights into tailocin–bacteria interactions, informing the potential use of tailocins in microbiome manipulation and antibacterial therapy.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic enhancer landscapes in human craniofacial development

The genetic basis of human facial variation and craniofacial birth defects remains poorly understood. Distant-acting transcriptional enhancers control the fine-tuned spatiotemporal expression of genes during critical stages of craniofacial development. However, a lack of accurate maps of the genomic locations and cell type-resolved activities of craniofacial enhancers prevents their systematic exploration in human genetics studies. Here, we combine histone modification, chromatin accessibility, and gene expression profiling of human craniofacial development with single-cell analyses of the developing mouse face to define the regulatory landscape of facial development at tissue- and single cell-resolution. We provide temporal activity profiles for 14,000 human developmental craniofacial enhancers. We find that 56% of human craniofacial enhancers share chromatin accessibility in the mouse and we provide cell population- and embryonic stage-resolved predictions of their in vivo activity. Taken together, our data provide an expansive resource for genetic and developmental studies of human craniofacial development.

60 APPLIED LIFE SCIENCES↗

Author Correction: Expanded encyclopaedias of DNA elements in the human and mouse genomes

In the version of this article initially published, two members of the ENCODE Project Consortium were missing from the author list. Rizi Ai (Department of Chemistry and Biochemistry, University of California, San Diego, La Jolla, CA, USA) and Shantao Li (Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA) are now included in the author list. These errors have been corrected in the online version of the article.

59 BASIC BIOLOGICAL SCIENCES↗

Long-read RNA sequencing atlas of human microglia isoforms elucidates disease-associated genetic regulation of splicing

Microglia, the innate immune cells of the central nervous system, have been genetically implicated in multiple neurodegenerative diseases. Mapping the genetics of gene expression in human microglia has identified several loci associated with disease-associated genetic variants in microglia-specific regulatory elements. However, identifying genetic effects on splicing is challenging because of the use of short sequencing reads. Here, we present the isoform-centric microglia genomic atlas (isoMiGA), which leverages long-read RNA sequencing to identify 35,879 novel microglia isoforms. We show that these isoforms are involved in stimulation response and brain region specificity. We then quantified the expression of both known and novel isoforms in a multi-ancestry meta-analysis of 555 human microglia short-read RNA sequencing samples from 391 donors, and found associations with genetic risk loci in Alzheimer’s and Parkinson’s disease. We nominate several loci that may act through complex changes in isoform and splice-site usage.

59 BASIC BIOLOGICAL SCIENCES↗