Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Defining the Antitumor Mechanism of Action of a Clinical-stage Compound as a Selective Degrader of the Nuclear Pore Complex

Cancer cells are acutely dependent on nuclear transport due to elevated transcriptional activity, suggesting an unrealized opportunity for selective therapeutic inhibition of the nuclear pore complex (NPC). Through large-scale phenotypic profiling of cancer cell lines, genome-scale functional genomic modifier screens, and mass spectrometry–based proteomics, we discovered that the clinical drug PRLX-93936 is a molecular glue that binds and reprograms the TRIM21 ubiquitin ligase to degrade the NPC. Upon compound-induced TRIM21 recruitment, the nuclear pore is ubiquitylated and degraded, resulting in the loss of short-lived cytoplasmic mRNA transcripts and the induction of cancer cell apoptosis. Direct compound binding to TRIM21 was confirmed via surface plasmon resonance and X-ray crystallography, whereas compound-induced TRIM21–nucleoporin complex formation was demonstrated through multiple orthogonal approaches in cells and in vitro. Phenotype-guided optimization yielded compounds with 10-fold greater potency and drug-like properties, along with robust pharmacokinetics and efficacy against pancreatic cancer xenografts and patient-derived organoids.

Yuan, Linjie [Stanford School of Medicine, CA (Uni↗

Genome-wide functional screens enable the prediction of high activity CRISPR-Cas9 and -Cas12a guides in Yarrowia lipolytica

Abstract Genome-wide functional genetic screens have been successful in discovering genotype-phenotype relationships and in engineering new phenotypes. While broadly applied in mammalian cell lines and in E. coli , use in non-conventional microorganisms has been limited, in part, due to the inability to accurately design high activity CRISPR guides in such species. Here, we develop an experimental-computational approach to sgRNA design that is specific to an organism of choice, in this case the oleaginous yeast Yarrowia lipolytica . A negative selection screen in the absence of non-homologous end-joining, the dominant DNA repair mechanism, was used to generate single guide RNA (sgRNA) activity profiles for both SpCas9 and LbCas12a. This genome-wide data served as input to a deep learning algorithm, DeepGuide, that is able to accurately predict guide activity. DeepGuide uses unsupervised learning to obtain a compressed representation of the genome, followed by supervised learning to map sgRNA sequence, genomic context, and epigenetic features with guide activity. Experimental validation, both genome-wide and with a subset of selected genes, confirms DeepGuide’s ability to accurately predict high activity sgRNAs. DeepGuide provides an organism specific predictor of CRISPR guide activity that with retraining could be applied to other fungal species, prokaryotes, and other non-conventional organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Workflow for High-throughput Screening of Enzyme Mutant Libraries Using Matrix-assisted Laser Desorption/Ionization Mass Spectrometry Analysis of Escherichia coli Colonies

High-throughput molecular screening of microbial colonies and DNA libraries are critical procedures that enable applications such as directed evolution, functional genomics, microbial identification, and creation of engineered microbial strains to produce high-value molecules. A promising chemical screening approach is the measurement of products directly from microbial colonies via optically guided matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). Measuring the compounds from microbial colonies bypasses liquid culture with a screen that takes approximately 5 s per sample. We describe a protocol combining a dedicated informatics pipeline and sample preparation method that can prepare up to 3,000 colonies in under 3 h. The screening protocol starts from colonies grown on Petri dishes and then transferred onto MALDI plates via imprinting. The target plate with the colonies is imaged by a flatbed scanner and the colonies are located via custom software. The target plate is coated with MALDI matrix, MALDI-MS analyzes the colony locations, and data analysis enables the determination of colonies with the desired biochemical properties. This workflow screens thousands of colonies per day without requiring additional automation. The wide chemical coverage and the high sensitivity of MALDI-MS enable diverse screening projects such as modifying enzymes and functional genomics surveys of gene activation/inhibition libraries.

Choe, Kisurb↗

16s Amplicon Analysis of Soil Data for Interactive effects of depth and differential irrigation on soil microbiome composition and functioning

Genomic DNA was isolated from soil and rhizosphere samples using the Zymo Quick-DNA fecal/soil microbe miniprep kit (catalog no. D6010) according to the manufacturer’s instructions (Zymo Research; Irvine, CA) with the modification of eluting in 100 uL elution buffer. Sample concentrations were quantified using the Qubit dsDNA HS assay kit (Thermo Fisher). For rhizosphere samples only, DNA was subsequently purified using Zymo’s ZR-96 DNA Clean & Concentrator kit (catalog no. D4024) to account for low DNA concentrations of these samples. . In each replicate block, there were five drip irrigation treatments (T1 = 100% normal irrigation, T2 = 56.25%, T3 = 37.5%, T4 = 18.75% and T5=no irrigation. On July 20, the strength of the drought treatments was increased: T1 remained at 100%, whereas T2 changed from 75% to 56.25%, T2 changed from 50% to 37.5%, and T4 changed from 25% to 18.75%. Sequencing was performed as described previously (Naylor, Fansler, et al. 2020). Sequences were amplified on the MiSeq platform (Illumina, San Diego, CA) using 16S primers (515F and 806R) specific to the V4 region. Raw sequence data was processed with the pipeline Hundo for amplicon quality control and annotation. Downstream statistical analyses on 16S datasets were performed using the program R and the packages ‘phyloseq’ and ‘vegan’.

Soil microbiome, metatranscriptomics↗

METABOLIC: high-throughput profiling of microbial genomes for functional traits, metabolism, biogeochemistry, and community-scale functional networks

Background Advances in microbiome science are being driven in large part due to our ability to study and infer microbial ecology from genomes reconstructed from mixed microbial communities using metagenomics and single-cell genomics. Such omics-based techniques allow us to read genomic blueprints of microorganisms, decipher their functional capacities and activities, and reconstruct their roles in biogeochemical processes. Currently available tools for analyses of genomic data can annotate and depict metabolic functions to some extent; however, no standardized approaches are currently available for the comprehensive characterization of metabolic predictions, metabolite exchanges, microbial interactions, and microbial contributions to biogeochemical cycling. Results We present METABOLIC (METabolic And BiogeOchemistry anaLyses In miCrobes), a scalable software to advance microbial ecology and biogeochemistry studies using genomes at the resolution of individual organisms and/or microbial communities. The genome-scale workflow includes annotation of microbial genomes, motif validation of biochemically validated conserved protein residues, metabolic pathway analyses, and calculation of contributions to individual biogeochemical transformations and cycles. The community-scale workflow supplements genome-scale analyses with determination of genome abundance in the microbiome, potential microbial metabolic handoffs and metabolite exchange, reconstruction of functional networks, and determination of microbial contributions to biogeochemical cycles. METABOLIC can take input genomes from isolates, metagenome-assembled genomes, or single-cell genomes. Results are presented in the form of tables for metabolism and a variety of visualizations including biogeochemical cycling potential, representation of sequential metabolic transformations, community-scale microbial functional networks using a newly defined metric “MW-score” (metabolic weight score), and metabolic Sankey diagrams. METABOLIC takes ~ 3 h with 40 CPU threads to process ~ 100 genomes and corresponding metagenomic reads within which the most compute-demanding part of hmmsearch takes ~ 45 min, while it takes ~ 5 h to complete hmmsearch for ~ 3600 genomes. Tests of accuracy, robustness, and consistency suggest METABOLIC provides better performance compared to other software and online servers. To highlight the utility and versatility of METABOLIC, we demonstrate its capabilities on diverse metagenomic datasets from the marine subsurface, terrestrial subsurface, meadow soil, deep sea, freshwater lakes, wastewater, and the human gut. Conclusion METABOLIC enables the consistent and reproducible study of microbial community ecology and biogeochemistry using a foundation of genome-informed microbial metabolism, and will advance the integration of uncultivated organisms into metabolic and biogeochemical models. METABOLIC is written in Perl and R and is freely available under GPLv3 at https://github.com/AnantharamanLab/METABOLIC.

59 BASIC BIOLOGICAL SCIENCES↗

Prokaryotic and eukaryotic cell-free systems for prototyping (CRADA Final Report)

CRADA FP00008491 between Berkeley Lab and Synvitrobio, Inc. (now Tierra Biosciences) validated the use of cell-free phenotyping to conduct functional genomics. Currently, most phenotyping work occurs using cellular fermentation methods and cellular techniques. This limits functional genomics throughput, which increasingly cannot handle the wealth of genetic information developed from next-generation sequencing technologies. De-risking a cell-free phenotyping approach has the advantage of increasing multiple-fold the throughput of genetic information that can be explored and expanding the $3B market for protein synthesis and characterization, leading to the accelerated development of human therapeutics and new biologically based materials.

59 BASIC BIOLOGICAL SCIENCES↗

Expanded genome and proteome reallocation in a novel, robust Bacillus coagulans strain capable of utilizing pentose and hexose sugars

Bacillus coagulans, a Gram-positive thermophilic bacterium, is recognized for its probiotic properties and recent development as a microbial cell factory. Despite its importance for biotechnological applications, the current understanding of B. coagulans’ robustness is limited, especially for undomesticated strains. To fill this knowledge gap, we characterized the metabolic capability and performed functional genomics and systems analysis of a novel, robust strain, B. coagulans B-768. Genome sequencing revealed that B-768 has the largest B. coagulans genome known to date (3.94 Mbp), about 0.63 Mbp larger than the average genome of sequenced B. coagulans strains, with expanded carbohydrate metabolism and mobilome. Functional genomics identified a well-equipped genetic portfolio for utilizing a wide range of C5 (xylose, arabinose), C6 (glucose, mannose, galactose), and C12 (cellobiose) sugars present in biomass hydrolysates, which was validated experimentally. For growth on individual xylose and glucose, the dominant sugars in biomass hydrolysates, B-768 exhibited distinct phenotypes and proteome profiles. Faster growth and glucose uptake rates resulted in lactate overflow metabolism, which makes B. coagulans a lactate overproducer; however, slower growth and xylose uptake diminished overflow metabolism due to the high energy demand for sugar assimilation. Carbohydrate Transport and Metabolism (COG-G), Translation (COG-J), and Energy Conversion and Production (COG-C) made up 60%–65% of the measured proteomes but were allocated differently when growing on xylose and glucose. The trade-off in proteome reallocation, with high investment in COG-C over COG-G, explains the xylose growth phenotype with significant upregulation of xylose metabolism, pyruvate metabolism, and tricarboxylic acid (TCA) cycle. Strain B-768 tolerates and effectively utilizes inhibitory biomass hydrolysates containing mixed sugars and exhibits hierarchical sugar utilization with glucose as the preferential substrate.

carbohydrate metabolism↗

A large sequenced mutant library – valuable reverse genetic resource that covers 98% of sorghum genes

SUMMARY Mutant populations are crucial for functional genomics and discovering novel traits for crop breeding. Sorghum , a drought and heat‐tolerant C4 species, requires a vast, large‐scale, annotated, and sequenced mutant resource to enhance crop improvement through functional genomics research. Here, we report a sorghum large‐scale sequenced mutant population with 9.5 million ethyl methane sulfonate (EMS)‐induced mutations that covered 98% of sorghum's annotated genes using inbred line BTx623. Remarkably, a total of 610 320 mutations within the promoter and enhancer regions of 18 000 and 11 790 genes, respectively, can be leveraged for novel research of cis ‐regulatory elements. A comparison of the distribution of mutations in the large‐scale mutant library and sorghum association panel (SAP) provides insights into the influence of selection. EMS‐induced mutations appeared to be random across different regions of the genome without significant enrichment in different sections of a gene, including the 5′ UTR, gene body, and 3′‐UTR. In contrast, there were low variation density in the coding and UTR regions in the SAP. Based on the K a / K s value, the mutant library (~1) experienced little selection, unlike the SAP (0.40), which has been strongly selected through breeding. All mutation data are publicly searchable through SorbMutDB ( https://www.depts.ttu.edu/igcast/sorbmutdb.php ) and SorghumBase ( https://sorghumbase.org/ ). This current large‐scale sequence‐indexed sorghum mutant population is a crucial resource that enriched the sorghum gene pool with novel diversity and a highly valuable tool for the Poaceae family, that will advance plant biology research and crop breeding.

59 BASIC BIOLOGICAL SCIENCES↗

Epigenetic effects associated with salmonid supplementation and domestication

Abstract Several studies have demonstrated lower fitness of salmonids born and reared in a hatchery setting compared to those born in nature, yet broad-scale genome-wide genetic differences between hatchery-origin and natural-origin fish have remained largely undetected. Recent research efforts have focused on using epigenetic tools to explore the role of heritable changes outside of genetic variation in response to hatchery rearing. We synthesized the results from salmonid studies that have directly compared methylation differences between hatchery-origin and natural-origin fish. Overall, the majority of studies found substantial differences in methylation patterns and overlap in functional genomic regions between hatchery-origin and natural-origin fish which have been replicated in parallel across geographical locations. Epigenetic differences were consistently found in the sperm of hatchery-origin versus natural-origin fish along with evidence for maternal effects, providing a potential source of multigenerational transmission. While there were clear epigenetic differences in gametic lines between hatchery-origin and natural-origin fish, only a limited number explored the potential mechanisms explaining these differences. We outline opportunities for epigenetics to inform salmonid breeding and rearing practices and to mitigate for fitness differences between hatchery-origin and natural-origin fish. We then provide possible explanations and avenues of future epigenetics research in salmonid supplementation programs, including: 1) further exploration of the factors in early development shaping epigenetic differences, 2) understanding the functional genomic changes that are occurring in response to epigenetic changes, 3) elucidating the relationship between epigenetics, phenotypic variation, and fitness, and 4) determining heritability of epigenetic marks along with persistence of marks across generations.

Koch, Ilana J. (ORCID:0000000335943108)↗

Protoplast-Based Transient Expression and Gene Editing in Shrub Willow ( Salix purpurea L .)

Shrub willows (Salix section Vetrix) are grown as a bioenergy crop in multiple countries and as ornamentals across the northern hemisphere. To facilitate the breeding and genetic advancement of shrub willow, there is a strong interest in the characterization and functional validation of genes involved in plant growth and biomass production. While protocols for shoot regeneration in tissue culture and production of stably transformed lines have greatly advanced this research in the closely related genus Populus, a lack of efficient methods for regeneration and transformation has stymied similar advancements in willow functional genomics. Moreover, transient expression assays in willow have been limited to callus tissue and hairy root systems. Here we report an efficient method for protoplast isolation from S. purpurea leaf tissue, along with transient overexpression and CRISPR-Cas9 mediated mutations. This is the first such report of transient gene expression in Salix protoplasts as well as the first application of CRISPR technology in this genus. These new capabilities pave the way for future functional genomics studies in this important bioenergy and ornamental crop.

59 BASIC BIOLOGICAL SCIENCES↗

Functional and genomic diversity of the sorghum phyllosphere microbiome

A collection of 47 bacteria isolated from the mucilage of aerial roots of energy sorghum is available at the Great Lakes Bioenergy Research Center, Michigan State University, Michigan, USA. We enriched bacteria with putative plant-beneficial phenotypes and included information on phenotypic diversity, taxonomy, and whole genome sequences.

agroecosystems↗

A multi-ancestry genetic study of pain intensity in 598,339 veterans

Chronic pain is a common problem, with more than one-fifth of adult Americans reporting pain daily or on most days. It adversely affects the quality of life and imposes substantial personal and economic costs. Efforts to treat chronic pain using opioids had a central role in precipitating the opioid crisis. Despite an estimated heritability of 25–50%, the genetic architecture of chronic pain is not well-characterized, in part because studies have largely been limited to samples of European ancestry. To help address this knowledge gap, we conducted a cross-ancestry meta-analysis of pain intensity in 598,339 participants in the Million Veteran Program, which identified 126 independent genetic loci, 69 of which are new. Pain intensity was genetically correlated with other pain phenotypes, level of substance use and substance use disorders, other psychiatric traits, education level and cognitive traits. Integration of the genome-wide association studies findings with functional genomics data shows enrichment for putatively causal genes (n = 142) and proteins (n = 14) expressed in brain tissues, specifically in GABAergic neurons. Drug repurposing analysis identified anticonvulsants, β-blockers and calcium-channel blockers, among other drug groups, as having potential analgesic effects. Our results provide insights into key molecular contributors to the experience of pain and highlight attractive drug targets.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440

ABSTRACT There is growing interest in engineering Pseudomonas putida KT2440 as a microbial chassis for the conversion of renewable and waste-based feedstocks, and metabolic engineering of P. putida relies on the understanding of the functional relationships between genes. In this work, independent component analysis (ICA) was applied to a compendium of existing fitness data from randomly barcoded transposon insertion sequencing (RB-TnSeq) of P. putida KT2440 grown in 179 unique experimental conditions. ICA identified 84 independent groups of genes, which we call fModules (“functional modules”), where gene members displayed shared functional influence in a specific cellular process. This machine learning-based approach both successfully recapitulated previously characterized functional relationships and established hitherto unknown associations between genes. Selected gene members from fModules for hydroxycinnamate metabolism and stress resistance, acetyl coenzyme A assimilation, and nitrogen metabolism were validated with engineered mutants of P. putida . Additionally, functional gene clusters from ICA of RB-TnSeq data sets were compared with regulatory gene clusters from prior ICA of RNAseq data sets to draw connections between gene regulation and function. Because ICA profiles the functional role of several distinct gene networks simultaneously, it can reduce the time required to annotate gene function relative to manual curation of RB-TnSeq data sets. IMPORTANCE This study demonstrates a rapid, automated approach for elucidating functional modules within complex genetic networks. While Pseudomonas putida randomly barcoded transposon insertion sequencing data were used as a proof of concept, this approach is applicable to any organism with existing functional genomics data sets and may serve as a useful tool for many valuable applications, such as guiding metabolic engineering efforts in other microbes or understanding functional relationships between virulence-associated genes in pathogenic microbes. Furthermore, this work demonstrates that comparison of data obtained from independent component analysis of transcriptomics and gene fitness datasets can elucidate regulatory-functional relationships between genes, which may have utility in a variety of applications, such as metabolic modeling, strain engineering, or identification of antimicrobial drug targets.

09 BIOMASS FUELS↗

DRAM example narrative

DRAM example narrative DRAM on KBase let's anyone run annotations using DRAM in the cloud. DRAM is an annotation tool that can annotate bacterial, archaeal and viral genomes and distills those annotatios into represetations of the functional genomic potential of those organisms. If you want to read more about DRAM you can check out the GitHub, wiki and journal article. DRAM annotate assemblies In KBase Assembly objects contain nucleotide sequences from genomes or metagenomes. DRAM can predict genes and annotate their function from KBase Assembly objects which may be microbial isolate genomes, metagenome assembled genomes or metagenomes. This is done with the Annotate and Distill Assemblies with DRAM app. This app can also anntoate AssemblySet objects which contain collection of Assembly objects. It also generates a Genome object and a GenomeSet object which can be used for further analysis with other KBase apps. The full annotations and other DRAM files are also available for download in the app.

59 BASIC BIOLOGICAL SCIENCES↗

A substrate-multiplexed platform for profiling enzymatic potential of plant family 1 glycosyltransferases

Plants have expanded various biosynthetic enzyme families to produce a wide diversity of natural products; however, most enzymes encoded in plant genomes remain uncharacterized, highlighting the need for new functional genomic approaches. Here, we report a platform enabling the rapid functional characterization of plant family 1 glycosyltransferases, which serve important roles in plant development, defense, and communication. Using substrate-multiplexed reactions, mass spectrometry, and automated analysis, we screen 85 enzymes against a diverse library of 453 natural products, for a total of nearly 40,000 possible reactions. The resulting dataset reveals a widespread promiscuity and a strong preference for planar, hydroxylated aromatic substrates among family 1 glycosyltransferases. We also characterize glycosyltransferases with an unusually wide substrate scope and with a non-canonical Cys-Asp catalytic dyad. This work establishes a widely-applicable enzymatic screening pipeline, reflects the immense glycosylation capability of plants, and has implications in biocatalysis, metabolic engineering, and gene discovery.

Sirirungruang, Sasilada↗

Validation of a metabolite–GWAS network for Populus trichocarpa family 1 UDP-glycosyltransferases

Metabolite genome-wide association studies (mGWASs) are increasingly used to discover the genetic basis of target phenotypes in plants such as Populus trichocarpa , a biofuel feedstock and model woody plant species. Despite their growing importance in plant genetics and metabolomics, few mGWASs are experimentally validated. Here, we present a functional genomics workflow for validating mGWAS-predicted enzyme–substrate relationships. We focus on uridine diphosphate–glycosyltransferases (UGTs), a large family of enzymes that catalyze sugar transfer to a variety of plant secondary metabolites involved in defense, signaling, and lignification. Glycosylation influences physiological roles, localization within cells and tissues, and metabolic fates of these metabolites. UGTs have substantially expanded in P. trichocarpa , presenting a challenge for large-scale characterization. Using a high-throughput assay, we produced substrate acceptance profiles for 40 previously uncharacterized candidate enzymes. Assays confirmed 10 of 13 leaf mGWAS associations, and a focused metabolite screen demonstrated varying levels of substrate specificity among UGTs. A substrate binding model case study of UGT-23 rationalized observed enzyme activities and mGWAS associations, including glycosylation of trichocarpinene to produce trichocarpin, a major higher-order salicylate in P. trichocarpa. We identified UGTs putatively involved in lignan, flavonoid, salicylate, and phytohormone metabolism, with potential implications for cell wall biosynthesis, nitrogen uptake, and biotic and abiotic stress response that determine sustainable biomass crop production. Our results provide new support for in silico analyses and evidence-based guidance for in vivo functional characterization.

59 BASIC BIOLOGICAL SCIENCES↗

Assembly, comparative analysis, and utilization of a single haplotype reference genome for soybean

Cultivar Williams 82 has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as Wm82-ISU-01 (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48 387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of Williams 82 reveal large regions of genomic heterogeneity, including regions of differential introgression from the cultivar Kingwa within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (> 20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the ‘Williams’ haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of Williams 82. In addition to the Wm82.a6 assembly, we also assembled the genome of ‘Fiskeby III,’ a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with Fiskeby III revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and Fiskeby III genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.

59 BASIC BIOLOGICAL SCIENCES↗