Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Functional characterization of prokaryotic dark matter: the road so far and what lies ahead

Eight-hundred thousand to one trillion prokaryotic species may inhabit our planet. Yet, fewer than two-hundred thousand prokaryotic species have been described. This uncharted fraction of microbial diversity, and its undisclosed coding potential, is known as the “microbial dark matter” (MDM). Next-generation sequencing has allowed to collect a massive amount of genome sequence data, leading to unprecedented advances in the field of genomics. Still, harnessing new functional information from the genomes of uncultured prokaryotes is often limited by standard classification methods. These methods often rely on sequence similarity searches against reference genomes from cultured species. This hinders the discovery of unique genetic elements that are missing from the cultivated realm. It also contributes to the accumulation of prokaryotic gene products of unknown function among public sequence data repositories, highlighting the need for new approaches for sequencing data analysis and classification. Increasing evidence indicates that these proteins of unknown function might be a treasure trove of biotechnological potential. Here, we outline the challenges, opportunities, and the potential hidden within the functional dark matter (FDM) of prokaryotes. We also discuss the pitfalls surrounding molecular and computational approaches currently used to probe these uncharted waters, and discuss future opportunities for research and applications.

59 BASIC BIOLOGICAL SCIENCES↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

Identifying microbial functional guilds performing cryptic organotrophic and lithotrophic redox cycles in anaerobic granular biofilms

Granular biofilms used in anaerobic digester systems contain diverse microbial populations that interact to hydrolyze organic matter and produce methane within controlled environments. Prior research investigated the feasibility of utilizing granular biofilms obtained from an anaerobic digester to remove nitrate without the addition of exogenous electron donors. These granules possessed a unique structure of alternating light and dark iron sulfide and pyrite rich layers that potentially served as both an electron source and sink, linking carbon, nitrogen, sulfur, and iron cycles. To characterize the functional roles of diverse microbial populations enriched within these layered biofilms, we analyzed metagenomes obtained from three different granules. Comparisons between the functional gene content of forty metagenome assembled genomes (MAGs) identified phylogenetically cohesive functional guilds. Each of these functional MAG clusters was assigned to specific steps in anaerobic digestion (hydrolysis, acidogenesis, acetogenesis, and methanogenesis) and anaerobic respiration (denitrification and sulfate reduction). Comparisons with metagenomes derived from a variety of natural and engineered ecosystems confirmed that the enriched denitrifying bacteria were similar to populations typically found in wetlands and biological nitrogen removal systems. Analysis of read alignments to individual genes within the forty MAGs identified conserved genomic features that were representative of the functions that distinguished functional guilds. Overall, this research illustrates the utility of functional based classification of microorganisms for characterizing ecosystem functions and highlights the potential application of engineered ecosystems to serve as experimental models for complex natural ecosystems.

Ecosystem engineering↗

CRISPR-GRIT: Guide RNAs with Integrated Repair Templates Enable Precise Multiplexed Genome Editing in the Diploid Fungal Pathogen Candida albicans

Candida albicans, an opportunistic fungal pathogen, causes severe infections in immunocompromised individuals. Limited classes and overuse of current antifungals have led to the rapid emergence of antifungal resistance. Thus, there is an urgent need to understand fungal pathogen genetics to develop new antifungal strategies. Genetic manipulation of C. albicans is encumbered by its diploid chromosomes requiring editing both alleles to elucidate gene function. Although the recent development of CRISPR-Cas systems has facilitated genome editing in C. albicans, large-scale and multiplexed functional genomic studies are still hindered by the necessity of cotransforming repair templates for homozygous knockouts. Here, we present CRISPR-GRIT (Guide RNAs with Integrated Repair Templates), a repair template-integrated guide RNA design for expedited gene knockouts and multiplexed gene editing in C. albicans. Here, we envision that this method can be used for high-throughput library screens and identification of synthetic lethal pairs in both C. albicans and other diploid organisms with strong homologous recombination machinery.

60 APPLIED LIFE SCIENCES↗

Shed Light in the DaRk LineagES of the Fungal Tree of Life—STRES

The polyphyletic group of black fungi within the Ascomycota (Arthoniomycetes, Dothideomycetes, and Eurotiomycetes) is ubiquitous in natural and anthropogenic habitats. Partly because of their dark, melanin-based pigmentation, black fungi are resistant to stresses including UV- and ionizing-radiation, heat and desiccation, toxic metals, and organic pollutants. Consequently, they are amongst the most stunning extremophiles and poly-extreme-tolerant organisms on Earth. Even though ca. 60 black fungal genomes have been sequenced to date, [mostly in the family Herpotrichiellaceae (Eurotiomycetes)], the class Dothideomycetes that hosts the largest majority of extremophiles has only been sparsely sampled. By sequencing up to 92 species that will become reference genomes, the “Shed light in The daRk lineagES of the fungal tree of life” (STRES) project will cover a broad collection of black fungal diversity spread throughout the Fungal Tree of Life. Interestingly, the STRES project will focus on mostly unsampled genera that display different ecologies and life-styles (e.g., ant- and lichen-associated fungi, rock-inhabiting fungi, etc.). With a resequencing strategy of 10- to 15-fold depth coverage of up to ~550 strains, numerous new reference genomes will be established. To identify metabolites and functional processes, these new genomic resources will be enriched with metabolomics analyses coupled with transcriptomics experiments on selected species under various stress conditions (salinity, dryness, UV radiation, oligotrophy). The data acquired will serve as a reference and foundation for establishing an encyclopedic database for fungal metagenomics as well as the biology, evolution, and ecology of the fungi in extreme environments.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of BBDuk metagenomic read trimming and decontamination

Background Investigators using metagenomic sequencing to study their microbiomes are often provided data that has been trimmed and decontaminated or do it themselves without knowing the effect these procedures can have on their downstream analyses. Here we evaluated the impact that JGI trimming and decontamination procedures had on assembly and binning metrics, placement of metagenome assembled genomes into species trees, and functional profiles of metagenome-assembled genomes (MAGs) extracted from twenty three complex rhizosphere metagenomes. We also investigated how more aggressive trimming impacts these binning metrics. Results We found that JGI trimmed and decontamination of input reads had some significant impacts in assembly and binning metrics compared to raw reads, and that differences in placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. More aggressive trimming beyond those used by JGI were found to reduce MAG counts. Conclusions Mild trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing? However, mild trimming and decontamination of metagenomic reads with high quality scores is recommended for those who elect to do so.

59 BASIC BIOLOGICAL SCIENCES↗

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS↗

Structural basis of differential gene expression at eQTLs loci from high-resolution ensemble models of 3D single-cell chromatin conformations

Abstract Motivation Techniques such as high-throughput chromosome conformation capture (Hi-C) have provided a wealth of information on nucleus organization and genome important for understanding gene expression regulation. Genome-Wide Association Studies have identified numerous loci associated with complex traits. Expression quantitative trait loci (eQTL) studies have further linked the genetic variants to alteration in expression levels of associated target genes across individuals. However, the functional roles of many eQTLs in noncoding regions remain unclear. Current joint analyses of Hi-C and eQTLs data lack advanced computational tools, limiting what can be learned from these data. Results We developed a computational method for simultaneous analysis of Hi-C and eQTL data, capable of identifying a small set of nonrandom interactions from all Hi-C interactions. Using these nonrandom interactions, we reconstructed large ensembles (×105) of high-resolution single-cell 3D chromatin conformations with thorough sampling, accurately replicating Hi-C measurements. Our results revealed many-body interactions in chromatin conformation at the single-cell level within eQTL loci, providing a detailed view of how 3D chromatin structures form the physical foundation for gene regulation, including how genetic variants of eQTLs affect the expression of associated eGenes. Furthermore, our method can deconvolve chromatin heterogeneity and investigate the spatial associations of eQTLs and eGenes at subpopulation level, revealing their regulatory impacts on gene expression. Together, ensemble modeling of thoroughly sampled single-cell chromatin conformations combined with eQTL data, helps decipher how 3D chromatin structures provide the physical basis for gene regulation, expression control, and aid in understanding the overall structure-function relationships of genome organization. Availability and implementation It is available at https://github.com/uic-liang-lab/3DChromFolding-eQTL-Loci.

Du, Lin (ORCID:0009000289869812)↗

Constructing the Nitrogen Flux Maps (NFMs) of Plants

The main objectives of this project are to construct plant N flux maps (NFMs) from plant genomes and to determine functionality of AT enzymes and plant N metabolic network. To address this grand challenge, this project made use of rapidly growing numbers of plant genomes, high-throughput functional characterization platforms, and computational modeling to deduce both biochemical and systems level functionality of ATs and NFMs. The obtained NFMs will provide a novel framework to advance basic understanding of plant N metabolism and facilitate rational engineering of plants with high productivity even under limited N input.

59 BASIC BIOLOGICAL SCIENCES↗

A universal and constant rate of gene content change traces pangenome flux to LUCA

Abstract Prokaryotic genomes constantly undergo gene flux via lateral gene transfer, generating a pangenome structure consisting of a conserved core genome surrounded by a more variable accessory genome shell. Over time, flux generates change in genome content. Here, we measure and compare the rate of genome flux for 5655 prokaryotic genomes as a function of amino acid sequence divergence in 36 universally distributed proteins of the informational core (IC). We find a clock of gene content change. The long-term average rate of gene content flux is remarkably constant across all higher prokaryotic taxa sampled, whereby the size of the accessory genome—the proportion of the genome harboring gene content difference for genome pairs—varies across taxa. The proportion of species-level accessory genes per genome, varies from 0% (Chlamydia) to 30%–33% (Alphaproteobacteria, Gammaproteobacteria, and Clostridia). A clock-like rate of gene content change across all prokaryotic taxa sampled suggest that pangenome structure is a general feature of prokaryotic genomes and that it has been in existence since the divergence of bacteria and archaea.

Microbiology↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

Evolutionary analysis of the LORELEI gene family in plants reveals regulatory subfunctionalization

Abstract A signaling complex comprising members of the LORELEI (LRE)-LIKE GPI-anchored protein (LLG) and Catharanthus roseus RECEPTOR-LIKE KINASE 1-LIKE (CrRLK1L) families perceive RAPID ALKALINIZATION FACTOR (RALF) peptides and regulate growth, reproduction, immunity, and stress responses in Arabidopsis (Arabidopsis thaliana). Genes encoding these proteins are members of multigene families in most angiosperms and could generate thousands of signaling complex variants. However, the links between expansion of these gene families and the functional diversification of this critical signaling complex as well as the evolutionary factors underlying the maintenance of gene duplicates remain unknown. Here, we investigated LLG gene family evolution by sampling land plant genomes and explored the function and expression of angiosperm LLGs. We found that LLG diversity within major land plant lineages is primarily due to lineage-specific duplication events, and that these duplications occurred both early in the history of these lineages and more recently. Our complementation and expression analyses showed that expression divergence (i.e. regulatory subfunctionalization), rather than functional divergence, explains the retention of LLG paralogs. Interestingly, all but one monocot and all eudicot species examined had an LLG copy with preferential expression in male reproductive tissues, while the other duplicate copies showed highest levels of expression in female or vegetative tissues. The single LLG copy in Amborella trichopoda is expressed vastly higher in male compared to in female reproductive or vegetative tissues. We propose that expression divergence plays an important role in retention of LLG duplicates in angiosperms.

Plant Sciences↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

Identification of over ten thousand candidate structured RNAs in viruses and phages

Structured RNAs play crucial roles in viruses, exerting influence over both viral and host gene expression. However, the extensive diversity of structured RNAs and their ability to act in cis or trans positions pose challenges for predicting and assigning their functions. While comparative genomics approaches have successfully predicted candidate structured RNAs in microbes on a large scale, similar efforts for viruses have been lacking. In this study, we screened over 5 million DNA and RNA viral sequences, resulting in the prediction of 10,006 novel candidate structured RNAs. These predictions are widely distributed across taxonomy and ecosystem. We found transcriptional evidence for 206 of these candidate structured RNAs in the human fecal microbiome. These candidate RNAs exhibited evidence of nucleotide covariation, indicative of selective pressure maintaining the predicted secondary structures. Our analysis revealed a diverse repertoire of candidate structured RNAs, encompassing a substantial number of putative tRNAs or tRNA-like structures, Rho-independent transcription terminators, and potentially cis-regulatory structures consistently positioned upstream of genes. In summary, our findings shed light on the extensive diversity of structured RNAs in viruses, offering a valuable resource for further investigations into their functional roles and implications in viral gene expression and pave the way for a deeper understanding of the intricate interplay between viruses and their hosts at the molecular level.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-omics analysis reveals the dynamic interplay between Vero host chromatin structure and function during vaccinia virus infection

The genome folds into complex configurations and structures thought to profoundly impact its function. The intricacies of this dynamic structure-function relationship are not well understood particularly in the context of viral infection. To unravel this interplay, here we provide a comprehensive investigation of simultaneous host chromatin structural (via Hi-C and ATAC-seq) and functional changes (via RNA-seq) in response to vaccinia virus infection. Over time, infection significantly impacts global and local chromatin structure by increasing long-range intra-chromosomal interactions and B compartmentalization and by decreasing chromatin accessibility and inter-chromosomal interactions. Local accessibility changes are independent of broad-scale chromatin compartment exchange (~12% of the genome), underscoring potential independent mechanisms for global and local chromatin reorganization. While infection structurally condenses the host genome, there is nearly equal bidirectional differential gene expression. Despite global weakening of intra-TAD interactions, functional changes including downregulated immunity genes are associated with alterations in local accessibility and loop domain restructuring. Therefore, chromatin accessibility and local structure profiling provide impactful predictions for host responses and may improve development of efficacious anti-viral counter measures including the optimization of vaccine design.

59 BASIC BIOLOGICAL SCIENCES↗

Genome and time-of-day transcriptome of Wolffia australiana link morphological minimization with gene loss and less growth control

Rootless plants in the genus Wolffia are some of the fastest growing known plants on Earth. Wolffia have a reduced body plan, primarily multiplying through a budding type of asexual reproduction. Here, we generated draft reference genomes for Wolffia australiana (Benth.) Hartog & Plas, which has the smallest genome size in the genus at 357 Mb and has a reduced set of predicted protein-coding genes at about 15,000. Comparison between multiple high-quality draft genome sequences from W. australiana clones confirmed loss of several hundred genes that are highly conserved among flowering plants, including genes involved in root developmental and light signaling pathways. Wolffia has also lost most of the conserved nucleotide-binding leucine-rich repeat (NLR) genes that are known to be involved in innate immunity, as well as those involved in terpene biosynthesis, while having a significant overrepresentation of genes in the sphingolipid pathways that may signify an alternative defense system. Diurnal expression analysis revealed that only 13% of Wolffia genes are expressed in a time-of-day (TOD) fashion, which is less than the typical ~40% found in several model plants under the same condition. In contrast to the model plants Arabidopsis and rice, many of the pathways associated with multicellular and developmental processes are not under TOD control in W. australiana, where genes that cycle the conditions tested predominantly have carbon processing and chloroplast-related functions. The Wolffia genome and TOD expression data set thus provide insight into the interplay between a streamlined plant body plan and optimized growth.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of key steps in the evolution of anaerobic methanotrophy in Candidatus Methanovorans (ANME-3) archaea

Despite their large environmental impact and multiple independent emergences, the processes leading to the evolution of anaerobic methanotrophic archaea (ANME) remain unclear. This work uses comparative metagenomics of a recently evolved but understudied ANME group, “Candidatus Methanovorans” (ANME-3), to identify evolutionary processes and innovations at work in ANME, which may be obscured in earlier evolved lineages. We identified horizontal transfer of hdrA homologs and convergent evolution in carbon and energy metabolic genes as potential early steps in Methanovorans evolution. We also identified the erosion of genes required for methylotrophic methanogenesis along with horizontal acquisition of multiheme cytochromes and other loci uniquely associated with ANME. The assembly and comparative analysis of multiple Methanovorans genomes offers important functional context for understanding the niche-defining metabolic differences between methane-oxidizing ANME and their methanogen relatives. Furthermore, this work illustrates the multiple evolutionary modes at play in the transition to a globally important metabolic niche.

59 BASIC BIOLOGICAL SCIENCES↗

Seagrass genomes reveal ancient polyploidy and adaptations to the marine environment

Here, we present chromosome-level genome assemblies from representative species of three independently evolved seagrass lineages: Posidonia oceanica, Cymodocea nodosa, Thalassia testudinum and Zostera marina. We also include a draft genome of Potamogeton acutifolius, belonging to a freshwater sister lineage to Zosteraceae. All seagrass species share an ancient whole-genome triplication, while additional whole-genome duplications were uncovered for C. nodosa, Z. marina and P. acutifolius. Comparative analysis of selected gene families suggests that the transition from submerged-freshwater to submerged-marine environments mainly involved fine-tuning of multiple processes (such as osmoregulation, salinity, light capture, carbon acquisition and temperature) that all had to happen in parallel, probably explaining why adaptation to a marine lifestyle has been exceedingly rare. Major gene losses related to stomata, volatiles, defence and lignification are probably a consequence of the return to the sea rather than the cause of it. These new genomes will accelerate functional studies and solutions, as continuing losses of the savannahs of the sea are of major concern in times of climate change and loss of biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗