Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Predicting metabolic modules in incomplete bacterial genomes with MetaPathPredict

The reconstruction of complete microbial metabolic pathways using ‘omics data from environmental samples remains challenging. Computational pipelines for pathway reconstruction that utilize machine learning methods to predict the presence or absence of KEGG modules in incomplete genomes are lacking. Here, we present MetaPathPredict, a software tool that incorporates machine learning models to predict the presence of complete KEGG modules within bacterial genomic datasets. Using gene annotation data and information from the KEGG module database, MetaPathPredict employs deep learning models to predict the presence of KEGG modules in a genome. MetaPathPredict can be used as a command line tool or as a Python module, and both options are designed to be run locally or on a compute cluster. Benchmarks show that MetaPathPredict makes robust predictions of KEGG module presence within highly incomplete genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Protoplast-Based Transient Expression and Gene Editing in Shrub Willow ( Salix purpurea L .)

Shrub willows (Salix section Vetrix) are grown as a bioenergy crop in multiple countries and as ornamentals across the northern hemisphere. To facilitate the breeding and genetic advancement of shrub willow, there is a strong interest in the characterization and functional validation of genes involved in plant growth and biomass production. While protocols for shoot regeneration in tissue culture and production of stably transformed lines have greatly advanced this research in the closely related genus Populus, a lack of efficient methods for regeneration and transformation has stymied similar advancements in willow functional genomics. Moreover, transient expression assays in willow have been limited to callus tissue and hairy root systems. Here we report an efficient method for protoplast isolation from S. purpurea leaf tissue, along with transient overexpression and CRISPR-Cas9 mediated mutations. This is the first such report of transient gene expression in Salix protoplasts as well as the first application of CRISPR technology in this genus. These new capabilities pave the way for future functional genomics studies in this important bioenergy and ornamental crop.

59 BASIC BIOLOGICAL SCIENCES↗

Marker-free carotenoid-enriched rice generated through targeted gene insertion using CRISPR-Cas9

Targeted insertion of transgenes at pre-determined plant genomic safe harbors provides a desirable alternative to insertions at random sites achieved through conventional methods. Most existing cases of targeted gene insertion in plants have either relied on the presence of a selectable marker gene in the insertion cassette or occurred at low frequency with relatively small DNA fragments (<1.8 kb). Here, we report the use of an optimized CRISPR-Cas9-based method to achieve the targeted insertion of a 5.2 kb carotenoid biosynthesis cassette at two genomic safe harbors in rice. We obtain marker-free rice plants with high carotenoid content in the seeds and no detectable penalty in morphology or yield. Whole-genome sequencing reveals the absence of off-target mutations by Cas9 in the engineered plants. These results demonstrate targeted gene insertion of marker-free DNA in rice using CRISPR-Cas9 genome editing, and offer a promising strategy for genetic improvement of rice and other crops.

59 BASIC BIOLOGICAL SCIENCES↗

pyFLANK, a graph neural network based null distribution inference model for F ST outlier detection

Detecting genomic regions under selection is essential for understanding how populations adapt to different environments, yet it remains challenging due to the confounding effects of demographic history and linkage disequilibrium (LD). Fixation index (F ST ) is a widely used statistic to identify genomic regions under adaptation. However, identifying genes under selection by defining F ST outliers often remains challenging, owing to confounding effects of underlying demographic history. Traditional methods assume independence among loci and rely on simple demographic models, while newer models perform much better but are computationally expensive and not easily scalable. Here, we present pyFLANK, an open-source and automated Python implementation which detects F ST outliers using a null distribution inferred from quasi-independent loci. Our tool integrates three approaches to identify loci obeying a null distribution: graph neural network (GNN) inference, linkage disequilibrium (LD)-based inference, and user-defined input. Because pyFLANK uses GNN-based inference of quasi-independent loci, it yields a more accurate null model with less need for user parameter input. In simulation experiments, pyFLANK achieved lower false positive rates than current methods while maintaining comparable detection power, indicating that its refined null model better distinguishes true adaptive loci from background variation. The GNN-based model, in particular, detected additional loci associated with phenotypic variance that were not identified by existing methods. Assessments of simulation and real data from different species demonstrate that pyFLANK achieves lower false positive rates compared with other commonly used F ST outlier detectors, while maintaining comparable detection power and excellent computational performance, providing a robust and user-friendly tool for identifying loci under divergent selection. It extends existing F ST outlier frameworks by incorporating explicit LD-aware strategies for null model calibration. The method is intended as a practical and scalable complement to existing genome scan approaches.

FST↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Improving CRISPR/Cas9-mediated genome editing efficiency in Yarrowia lipolytica using direct tRNA-sgRNA fusions

We report Yarrowia lipolytica is an important oleaginous yeast currently used in the production of specialty chemicals and has a great potential for further applications in lipid biotechnology. Harnessing the full potential of Y. lipolytica is, however, limited by its inherent recalcitrance to genetic manipulation. In contrast to Saccharomyces cerevisiae, Y. lipolytica is poor in homology-mediated DNA repair and thus in homologous recombination, which limits site-specific gene editing in this yeast. Recently developed CRISPR/Cas9-based methods using tRNA-sgRNA fusions succeeded in editing some genomic loci in Y. lipolytica. Nonetheless, the majority of other tested loci either failed editing or editing was achieved but at very low efficiency using these methods. Using tools of secondary RNA structure prediction, we were able to improve the design of the tRNA-sgRNA fusions used for the expression of single guide RNA (sgRNA) in such methods. This resulted in high efficiency CRISPR/cas9 gene editing at chromosomal loci that failed gene editing or were edited at very low efficiencies with previous methods. In addition, we characterized the gene editing performance of our newly designed tRNA-sgRNA fusions for both chromosomal gene integration and deletion. As such, this study presents an efficient CRISPR/Cas9-mediated gene-editing tool for efficient genetic engineering of Yarrowia lipolytica.

59 BASIC BIOLOGICAL SCIENCES↗

Data for Transposon Signatures of Allopolyploid Genome Evolution

Hybridization brings together chromosome sets from two or more distinct progenitor species. Genome duplication associated with hybridization, or allopolyploidy, allows these chromosome sets to persist as distinct subgenomes during subsequent meioses. Here, we present a general method for identifying the subgenomes of a polyploid based on shared ancestry as revealed by the genomic distribution of repetitive elements that were active in the progenitors. This subgenome-enriched transposable element signal is intrinsic to the polyploid, allowing broader applicability than other approaches that depend on the availability of sequenced diploid relatives. We develop the statistical basis of the method, demonstrate its applicability in the well-studied cases of tobacco, cotton, and Brassica napus, and apply it to several cases: allotetraploid cyprinids, allohexaploid false flax, and allooctoploid strawberry. These analyses provide insight into the origins of these polyploids, revise the subgenome identities of strawberry, and provide perspective on subgenome dominance in higher polyploids.

Genomics↗

Halophytes and heavy metals: A multi‐omics approach to understand the role of gene and genome duplication in the abiotic stress tolerance of Cakile maritima

Abstract Premise The origin of diversity is a fundamental biological question. Gene duplications are one mechanism that provides raw material for the emergence of novel traits, but evolutionary outcomes depend on which genes are retained and how they become functionalized. Yet, following different duplication types (polyploidy and tandem duplication), the events driving gene retention and functionalization remain poorly understood. Here we usedCakile maritima, a species that is tolerant to salt and heavy metals and shares an ancient whole‐genome triplication with closely related salt‐sensitive mustard crops (Brassica), as a model to explore the evolution of abiotic stress tolerance following polyploidy. Methods Using a combination of ionomics, free amino acid profiling, and comparative genomics, we characterize aspects of salt stress response inC. maritimaand identify retained duplicate genes that have likely enabled adaptation to salt and mild levels of cadmium. Results Cakile maritimais tolerant to both cadmium and salt treatments through uptake of cadmium in the roots. Proline constitutes greater than 30% of the free amino acid pool inC. maritimaand likely contributes to abiotic stress tolerance. We find duplicated gene families are enriched in metabolic and transport processes and identify key transport genes that may be involved inC. maritimaabiotic stress tolerance. Conclusions These findings identify pathways and genes that could be used to enhance plant resilience and provide a putative understanding of the roles of duplication types and retention on the evolution of abiotic stress response.

Plant Sciences↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

Targeted assemblies of cas1 suggest CRISPR-Cas’s response to soil warming

Abstract There is an increasing interest in the clustered regularly interspaced short palindromic repeats CRISPR-associated protein (CRISPR-Cas) system to reveal potential virus–host dynamics. The universal and most conserved Cas protein, cas1 is an ideal marker to elucidate CRISPR-Cas ecology. We constructed eight Hidden Markov Models (HMMs) and assembled cas1 directly from metagenomes by a targeted-gene assembler, Xander, to improve detection capacity and resolve the diverse CRISPR-Cas systems. The eight HMMs were first validated by recovering all 17 cas1 subtypes from the simulated metagenome generated from 91 prokaryotic genomes across 11 phyla. We challenged the targeted method with 48 metagenomes from a tallgrass prairie in Central Oklahoma recovering 3394 cas1. Among those, 88 were near full length, 5 times more than in de-novo assemblies from the Oklahoma metagenomes. To validate the host assignment by cas1, the targeted-assembled cas1 was mapped to the de-novo assembled contigs. All the phylum assignments of those mapped contigs were assigned independent of CRISPR-Cas genes on the same contigs and consistent with the host taxonomies predicted by the mapped cas1. We then investigated whether 8 years of soil warming altered cas1 prevalence within the communities. A shift in microbial abundances was observed during the year with the biggest temperature differential (mean 4.16 °C above ambient). cas1 prevalence increased and even in the phyla with decreased microbial abundances over the next 3 years, suggesting increasing virus–host interactions in response to soil warming. This targeted method provides an alternative means to effectively mine cas1 from metagenomes and uncover the host communities.

54 ENVIRONMENTAL SCIENCES↗

$\mathrm{CROPSR}$: an automated platform for complex genome-wide $\mathrm{CRISPR}$ g$\mathrm{RNA}$ design and validation

CRISPR/Cas9 technology has become an important tool to generate targeted, highly specific genome mutations. The technology has great potential for crop improvement, as crop genomes are tailored to optimize specific traits over generations of breeding. Many crops have highly complex and polyploid genomes, particularly those used for bioenergy or bioproducts. The majority of tools currently available for designing and evaluating gRNAs for CRISPR experiments were developed based on mammalian genomes that do not share the characteristics or design criteria for crop genomes. We have developed an open source tool for genome-wide design and evaluation of gRNA sequences for CRISPR experiments, CROPSR. The genome-wide approach provides a significant decrease in the time required to design a CRISPR experiment, including validation through PCR, at the expense of an overhead compute time required once per genome, at the first run. To better cater to the needs of crop geneticists, restrictions imposed by other packages on design and evaluation of gRNA sequences were lifted. A new machine learning model was developed to provide scores while avoiding situations in which the currently available tools sometimes failed to provide guides for repetitive, A/T-rich genomic regions. We show that our gRNA scoring model provides a significant increase in prediction accuracy over existing tools, even in non-crop genomes. CROPSR provides the scientific community with new methods and a new workflow for performing CRISPR/Cas9 knockout experiments. CROPSR reduces the challenges of working in crops, and helps speed gRNA sequence design, evaluation and validation. We hope that the new software will accelerate discovery and reduce the number of failed experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Abstract Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcript’s impacts on plant metabolism has employed data from controlled environments. Here, we generated and analyzed data on the variation in the abundance of 26 metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies (GWAS)) and transcripts (using both transcriptome-wide association studies (TWAS) and an explainable artificial intelligence (AI) approach based on random forest (RF)) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

Mathivanan, Ramesh Kanna↗

Evaluation of DNA Extraction Efficiency in Diverse Algae Strains Using Commercial Kits and Lysis Approaches

Efficient DNA extraction is essential for accurately monitoring microalgae communities in large-scale cultivation systems such as raceway ponds and wastewater ponds. Traditional phenol chloroform extracts are a staple in microbiology but are obsolete for routine sampling due to its high toxicity reagents and time intensive setups. Commercial DNA extraction kits are more favorable for the microbes found in these ponds, but lack specific kits made for these communities. Little is known about which kits perform the best, leading researchers to use a variety of different kits with inconsistent results. This project compared one precipitation based commercial kit (Lucigen Masterpure) and five wash based kits (Monarch, Zymo Quick-DNA, and three Qiagen DNeasy kits) using four brackish algae strains to determine which methods yield the greatest quantity and quality of genomic DNA. Extractions were evaluated using the manufacturers protocol, and additional pretreatment options were administered before a single kit to compare its potential in being added routinely before extractions. Pretreatment options included both cryogenic freeze-thawing and heat incubation using enzymes. DNA was quantified using Qubit fluorometry and NanoDrop purity ratios. Overall, the Qiagen PowerWater kit provided the highest DNA yield and purity, but at a significantly higher cost then the precipitation-based kit (MasterPure). It was also noted that while the precipitation-based kit was significantly cheaper, provided similar results, it took significantly more time to complete a single run. Cryogenic pretreatment (6x cycles) increased average DNA yields by up to 80%, whereas enzymatic pretreatment most improved purity ratios without substantially improving quantity. The results suggest that it may be more cost and time efficient to use Qiagen kits with the addition of lysis pretreatments to procure better results. Future works includes developing a better system to efficiently collect multi variable data, and to upscale to artificial polycultures using similar methodologies alongside sequencing to confirm kit results.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Genomics and Transcriptomics Analyses Reveal Divergent Plant Biomass-Degrading Strategies in Fungi

Plant biomass is one of the most abundant renewable carbon sources, which holds great potential for replacing current fossil-based production of fuels and chemicals. In nature, fungi can efficiently degrade plant polysaccharides by secreting a broad range of carbohydrate-active enzymes (CAZymes), such as cellulases, hemicellulases, and pectinases. Due to the crucial role of plant biomass-degrading (PBD) CAZymes in fungal growth and related biotechnology applications, investigation of their genomic diversity and transcriptional dynamics has attracted increasing attention. In this project, we systematically compared the genome content of PBD CAZymes in six taxonomically distant species, Aspergillus niger, Aspergillus nidulans, Penicillium subrubescens, Trichoderma reesei, Phanerochaete chrysosporium, and Dichomitus squalens, as well as their transcriptome profiles during growth on nine monosaccharides. Considerable genomic variation and remarkable transcriptomic diversity of CAZymes were identified, implying the preferred carbon source of these fungi and their different methods of transcription regulation. In addition, the specific carbon utilization ability inferred from genomics and transcriptomics was compared with fungal growth profiles on corresponding sugars, to improve our understanding of the conversion process. This study enhances our understanding of genomic and transcriptomic diversity of fungal plant polysaccharide-degrading enzymes and provides new insights into designing enzyme mixtures and metabolic engineering of fungi for related industrial applications.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid identification of enteric bacteria from whole genome sequences using average nucleotide identity metrics

Identification of enteric bacteria species by whole genome sequence (WGS) analysis requires a rapid and an easily standardized approach. We leveraged the principles of average nucleotide identity using MUMmer (ANIm) software, which calculates the percent bases aligned between two bacterial genomes and their corresponding ANI values, to set threshold values for determining species consistent with the conventional identification methods of known species. The performance of species identification was evaluated using two datasets: the Reference Genome Dataset v2 (RGDv2), consisting of 43 enteric genome assemblies representing 32 species, and the Test Genome Dataset (TGDv1), comprising 454 genome assemblies which is designed to represent all species needed to query for identification, as well as rare and closely related species. The RGDv2 contains six Campylobacter spp., three Escherichia/Shigella spp., one Grimontia hollisae, six Listeria spp., one Photobacterium damselae, two Salmonella spp., and thirteen Vibrio spp., while the TGDv1 contains 454 enteric bacterial genomes representing 42 different species. The analysis showed that, when a standard minimum of 70% genome bases alignment existed, the ANI threshold values determined for these species were ≥95 for Escherichia/Shigella and Vibrio species, ≥93% for Salmonella species, and ≥92% for Campylobacter and Listeria species. Using these metrics, the RGDv2 accurately classified all validation strains in TGDv1 at the species level, which is consistent with the classification based on previous gold standard methods.

59 BASIC BIOLOGICAL SCIENCES↗

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii V.2

The head smut (Tilletia maclaganii) is a significant pathogen of the bioenergy crop switchgrass. T. maclaganii typically is more prevalent in older stands of switchgrass and can contribute to significant biomass loss. Here, we outline the methods for the sequencing, assembly, and annotation of the first reference genome for Tilletia maclaganii.

Benucci, Gian Maria Niccolò [GLBRC - Michigan Stat↗

Barcoded reciprocal hemizygosity analysis via sequencing illuminates the complex genetic basis of yeast thermotolerance

Decades of successes in statistical genetics have revealed the molecular underpinnings of traits as they vary across individuals of a given species. But standard methods in the field cannot be applied to divergences between reproductively isolated taxa. Genome-wide reciprocal hemizygosity mapping (RH-seq), a mutagenesis screen in an interspecies hybrid background, holds promise as a method to accelerate the progress of interspecies genetics research. Here, we describe an improvement to RH-seq in which mutants harbor barcodes for cheap and straightforward sequencing after selection in a condition of interest. As a proof of concept for the new tool, we carried out genetic dissection of the difference in thermotolerance between two reproductively isolated budding yeast species. Experimental screening identified dozens of candidate loci at which variation between the species contributed to the thermotolerance trait. Hits were enriched for mitosis genes and other housekeeping factors, and among them were multiple loci with robust sequence signatures of positive selection. Together, these results shed new light on the mechanisms by which evolution solved the problems of cell survival and division at high temperature in the yeast clade, and they illustrate the power of the barcoded RH-seq approach.

59 BASIC BIOLOGICAL SCIENCES↗

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]↗