Engineering PapersSearch

SEARCH · Engineering Papers

Results for “genomic selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Data for Miscanthus giganteus Biolistic Transformation Using the Visible RUBY Red Marker Gene to Monitor Transformation Efficiency

Miscanthus × giganteus ( M × g ) is a high- yielding perennial C4 bioenergy crop, but genetic improvement by breeding is constrained by triploid sterility and clonal propagation. Improving genetic transformation methods for M × g would provide opportunities for advantageous trait introgression. Use of an easy to phenotype reporter gene is a promising strategy to improve transformation processes and efficiency. This study presents an efficient novel method for biolistic transformation of inflorescence- derived callus in M × g and demonstrates its efficacy using RUBY, a betalain-based noninvasive reporter that is visible throughout the transformation process. RUBY expression ( Zea mays codon optimized) was visible from callus stage through plantlet development into maturity. RUBY expressing independently transformed plants were confirmed by hygromycin phosphotransferase ELISA and by genomic PCR demonstrating that the RUBY phenotype is sufficient for screening transformants. The Zea mays codon optimized hygromycin selection marker was driven by previously established promoters for Miscanthus, ZmUBI and 2×35S, while the RUBY gene expression was controlled by a known Zea mays C4 promoter, Brachypodium UBI10, newly employed in Miscanthus. The construct containing the 2×35S promoter for hygromycin had a 15.1% transformation efficiency while the ZmUBI promoter had a 20.5% transformation efficiency. This study provides a novel, highly efficient protocol for successful biolistic transformation of M × g for stable expression. This study also demonstrates that RUBY expression can be used as a convenient and powerful monitor of transformation in ongoing and future work to engineer M × g into an improved bioproduct feedstock. **NOTE: in "TableS2_ProtocolComparison.csv", the data from row 665 to 971 should be removed.

Gene Editing

Optimizing genomic prediction for complex traits via investigating multiple factors in switchgrass

Genomic prediction has accelerated breeding processes and provided mechanistic insights into the genetic bases of complex traits. To further optimize genomic prediction, we assess the impact of genome assemblies, genotyping approaches, variant types, allelic complexities, polyploidy levels, and population structures on the prediction of 20 complex traits in switchgrass (Panicum virgatum L.), a perennial biofuel feedstock. Surprisingly, short read-based genome assembly performs comparably to or even better than long read-based assembly. Due to higher gene coverage, exome capture and multi-allelic variants outperform genotyping-by-sequencing and bi-allelic variants, respectively. Tetraploid models show higher prediction accuracy than octoploid models for most traits, likely due to the greater genetic distances among tetraploids. Depending on the trait in question, different types of variants need to be integrated for optimal predictions. Furthermore, our study provides insights into the factors influencing genomic prediction outcomes, guiding best practices for future studies and for improving agronomic traits in switchgrass and other species through selective breeding.

60 APPLIED LIFE SCIENCES

Beyond Component Optimization: Systems Level Biodesign for Lanthanide Recovery

Global demand for lanthanides (Ln) is projected to rise sharply over the next decade, while geographically concentrated supply chains and the low concentrations and matrix complexity of secondary feedstocks limit the reach of conventional hydro- and pyrometallurgical separation. Engineered biological systems offer a selective, low-energy alternative, and component-level advances in Ln-binding proteins, AI-designed selective scaffolds, and cell-surface display platforms now rival synthetic chelators in affinity and selectivity. These components, however, remain functionally isolated. Currently, there are no engineered chassis coupling recognition, intracellular trafficking, accumulation, and controlled release into an end-to-end pipeline. Here, we outline how new biodesign strategies and chassis selection must move beyond bioleaching to encompass the full recovery pathway. Achieving this requires integrating AI/ML-guided design, genome-scale build tools, high-throughput phenotyping, and biophysical transport modeling within a Design–Build–Test–Learn cycle tuned to recognition, trafficking, accumulation, and release.

Biodesign

Benchmarking Computational Tools for Calling SNPs and Indels in Complex Microbial Populations

The NASA BioNutrients missions seek to understand the suitability of microorganisms for bioproduction during space flight. One topic of interest is the stability of microbial genomes during long-term ambient storage and subsequent rehydration and growth. To address these questions, samples from 8 species were flown to ISS for 5 years of desiccated storage at ambient temperature (Stasis Packs) and 2 species were packaged along with powdered media inside a bioreactor system to allow hydration and growth in microgravity (Production Packs). For both systems, Whole Genome Sequencing (WGS) of the DNA extracted from the returned samples and paired ground controls will be conducted to identify changes in genome stability due to time, storage conditions and growth in space. Across the technical replicates, ground controls, 10 timepoints, and multiple experimental conditions, ~300 samples have been selected for initial analysis with WGS sequencing to 100x coverage. A flexible and resource efficient mutation calling pipeline is needed to process this large dataset and allow for comparisons between species. Many bioinformatics tools for calling Indels and Single Nucleotide Variants (SNVs) are designed for use with pure isolates, where true variations from the reference genome are expected to dominate the reads aligning to the location of mutation. In contrast, DNA from the Stasis Pack (SP) samples was collected directly after recovery from desiccated storage and the Production Pack (PP) samples were collected after fermentation. In this context, reads with mutations are expected to be less frequent than reads that align with the reference genome, as each sample will include multiple lines of cells. Thus, BioNutrients samples are expected to be similar to samples from cancer cell or “pooled” sequencing approaches. In preparation for the analysis of the BioNutrients samples, we have tested three mutation calling tools (GATK for Microbes, BreSeq and DiscoSNP) designed for complex samples. A challenge of validating mutation identification pipelines is a lack of “Ground Truth” datasets, especially for complex samples. To compare these three tools, we sought to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO as part of the Space Algae-1 mission. Here we present a summary of these tools against the analysis originally conducted using the CRISP tool. Critical metrics are compared such as runtime, the number of SNPs, the number and size of Indels, and patterns of transversion and transitions identified by each tool are reported. By sharing these benchmarking results collected in support of the BioNutrients mission, we aim to guide others seeking to identify SNVs in similarly complex microbial samples.

Biology

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES

Defining the Antitumor Mechanism of Action of a Clinical-stage Compound as a Selective Degrader of the Nuclear Pore Complex

Cancer cells are acutely dependent on nuclear transport due to elevated transcriptional activity, suggesting an unrealized opportunity for selective therapeutic inhibition of the nuclear pore complex (NPC). Through large-scale phenotypic profiling of cancer cell lines, genome-scale functional genomic modifier screens, and mass spectrometry–based proteomics, we discovered that the clinical drug PRLX-93936 is a molecular glue that binds and reprograms the TRIM21 ubiquitin ligase to degrade the NPC. Upon compound-induced TRIM21 recruitment, the nuclear pore is ubiquitylated and degraded, resulting in the loss of short-lived cytoplasmic mRNA transcripts and the induction of cancer cell apoptosis. Direct compound binding to TRIM21 was confirmed via surface plasmon resonance and X-ray crystallography, whereas compound-induced TRIM21–nucleoporin complex formation was demonstrated through multiple orthogonal approaches in cells and in vitro. Phenotype-guided optimization yielded compounds with 10-fold greater potency and drug-like properties, along with robust pharmacokinetics and efficacy against pancreatic cancer xenografts and patient-derived organoids.

Yuan, Linjie [Stanford School of Medicine, CA (Uni

PMA-Linked Fluorescence for Rapid Detection of Viable Bacterial Endospores

The most common approach for assessing the abundance of viable bacterial endospores is the culture-based plating method. However, culture-based approaches are heavily biased and oftentimes incompatible with upstream sample processing strategies, which make viable cells/spores uncultivable. This shortcoming highlights the need for rapid molecular diagnostic tools to assess more accurately the abundance of viable spacecraft-associated microbiota, perhaps most importantly bacterial endospores. Propidium monoazide (PMA) has received a great deal of attention due to its ability to differentiate live, viable bacterial cells from dead ones. PMA gains access to the DNA of dead cells through compromised membranes. Once inside the cell, it intercalates and eventually covalently bonds with the double-helix structures upon photoactivation with visible light. The covalently bound DNA is significantly altered, and unavailable to downstream molecular-based manipulations and analyses. Microbiological samples can be treated with appropriate concentrations of PMA and exposed to visible light prior to undergoing total genomic DNA extraction, resulting in an extract comprised solely of DNA arising from viable cells. This ability to extract DNA selectively from living cells is extremely powerful, and bears great relevance to many microbiological arenas.

LaDuc, Myron T.

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor

High Throughput Genome Releaser

In this study, we present the development of a High Throughput Genome Releaser, an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed for rapid, cost-effective, and efficient DNA extraction, optimized for subsequent PCR reactions. Our experimentation with various synthetic materials led us to select a particular type of plastic that mirrors the properties of glass cover slides, providing a smooth surface and effective compression capabilities. We engineered a 96-well device equipped with a 96-well plate and a top rod, operable both manually and automatically, which is compatible with widely used liquid-handling robot decks. This compatibility enhances ease of use in high-throughput PCR setups. Additionally, we developed software to support its automatic functions. The genome releaser facilitates the extraction of PCR-amplifiable genomic DNA from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells. This versatility could significantly advance biomanufacturing processes.

42 ENGINEERING

CAZyme domain architectures suggest fine-scale functional differentiation among anaerobic fungi and bacteria during lignocellulose conversion to volatile fatty acids

Anaerobic fermentation with microbial communities (microbiomes) is an emerging platform for conversion of lignocellulosic biomass to biofuels and bioproducts. The process relies on diverse anaerobic microbes that interact to deconstruct and convert lignocellulosic biomass into a range of products, such as volatile fatty acids (VFAs), which can be achieved by arresting methanogenesis during fermentation. However, defining the distinct functional roles played by various fungi and bacteria during anaerobic biodegradation remains poorly understood. Here, we performed parallel enrichment experiments from cow faeces, goat faeces, and anaerobic digester sludge, selecting for fungal or bacterial dominated communities that convert sorghum biomass into VFAs. Subsequently we reconstructed metabolic networks across these enrichments based on recovered bacterial metagenome-assembled genomes (MAGs) and fungal isolate genomes and profiled their metabolic activity using metatranscriptomics to identify potential functional niches. Our findings implicate diverse bacteria affiliated with the Bacteroidales and Lachnospiraceae in the direct conversion of lignocellulosic biomass to propionate and butyrate, respectively, whereas Neocallimastix-dominated fungal enrichments converted lignocellulose to lactate, acetate and formate. Analysis of carbohydrate-active enzymes (CAZymes) revealed fine-scale differences between microbes that expressed unique multi-functional enzymes linking two or more CAZymes together with distinct carbohydrate binding motifs, implicating lignocellulose structure as a key driver of selection and niche differentiation. Most of these multi-functional enzymes localized complementary degradation functions together, likely conferring synergistic degradation effects within and between microbiome members. We anticipate that these findings will help inform efforts to develop synthetic microbiomes with tailored functionality for low-cost conversion of lignocellulosic biomass to fuels and bio-based chemicals.

Lawson, Christopher E [University of Toronto;]

The Twins Study: NASA's First Foray into 21st Century Omics Research

The full array of 21st century omics-based research methods should be intelligently employed to reduce the health and performance risks that astronauts will be exposed to during exploration missions beyond low Earth Orbit. In March of 2015, US Astronaut Scott Kelly will launch to the International Space Station for a one year mission while his twin brother, Mark Kelly, a retired US Astronaut, remains on the ground. This situation presents an extremely rare flight opportunity to perform an integrated omics-based demonstration pilot study involving identical twin astronauts. A group of 10 principal investigators has been competitively selected, funded, and teamed together to form the Twins Study. A very broad range of biological function are being examined including the genome, epigenome, transcriptome, proteome, metabolome, gut microbiome, immunological response to vaccinations, indicators of atherosclerosis, physiological fluid shifts, and cognition. The plans for the Twins Study and an overview of initial results will be described as well as the technological and ethical issues raised for such spaceflight studies. An anticipated outcome of the Twins Study is that it will place NASA on a trajectory of using omics-based information to develop precision countermeasures for individual astronauts.

Kundrot, C. E.

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]

Phylogenetic and physiological diversity of microorganisms isolated from a deep greenland glacier ice core

We studied a sample from the GISP 2 (Greenland Ice Sheet Project) ice core to determine the diversity and survival of microorganisms trapped in the ice at least 120,000 years ago. Previously, we examined the phylogenetic relationships among 16S ribosomal DNA (rDNA) sequences in a clone library obtained by PCR amplification from genomic DNA extracted from anaerobic enrichments. Here we report the isolation of nearly 800 aerobic organisms that were grouped by morphology and amplified rDNA restriction analysis patterns to select isolates for further study. The phylogenetic analyses of 56 representative rDNA sequences showed that the isolates belonged to four major phylogenetic groups: the high-G+C gram-positives, low-G+C gram-positives, Proteobacteria, and the Cytophaga-Flavobacterium-Bacteroides group. The most abundant and diverse isolates were within the high-G+C gram-positive cluster that had not been represented in the clone library. The Jukes-Cantor evolutionary distance matrix results suggested that at least 7 isolates represent new species within characterized genera and that 49 are different strains of known species. The isolates were further categorized based on the isolation conditions, temperature range for growth, enzyme activity, antibiotic resistance, presence of plasmids, and strain-specific genomic variations. A significant observation with implications for the development of novel and more effective cultivation methods was that preliminary incubation in anaerobic and aerobic liquid prior to plating on agar media greatly increased the recovery of CFU from the ice core sample.

Bacteria/classification/genetics/isolation & purif

A ribozyme that triphosphorylates RNA 50-hydroxyl groups

The RNA world hypothesis describes a stage in the early evolution of life in which RNA served as genome and as the only genome-encoded catalyst. To test whether RNA world organisms could have used cyclic trimetaphosphate as an energy source, we developed an in vitro selection strategy for isolating ribozymes that catalyze the triphosphorylation of RNA 50 -hydroxyl groups with trimetaphosphate. Several active sequences were isolated, and one ribozyme was analyzed in more detail. The ribozyme was truncated to 96 nt, while retaining full activity. It was converted to a transformat and reacted with rates of 0.16 min1 under optimal conditions. The secondary structure appears to contain a four-helical junction motif. This study showed that ribozymes can use trimetaphosphate to triphosphorylate RNA 50 -hydroxyl groups and suggested that RNA world organisms could have used trimetaphosphate as their energy source.

Janina E. Moretti

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan

Development, optimization, and application of an episomal plasmid system for Rhodotorula toruloides

Rhodotorula toruloides is an emerging oleaginous yeast with strong potential as a microbial cell factory for the production of acetyl-CoA-derived bioproducts. However, engineering of this organism has been limited by the absence of a functional episomal plasmid system, a foundational genetic tool for rapid gene expression, pathway testing, and CRISPR-based genome engineering. Here, we report the first episomal plasmid system for R. toruloides . Through systematic screening of candidate autonomously replicating sequences (ARSs) from diverse sources, we identified multiple functional ARS elements and selected C63F4, a fragment derived from Contig 63 of R. toruloides CBS14, because of its stable performance. The resulting pC63F4 plasmid was maintained episomally, supported GFP reporter expression, exhibited a copy number of 2.39 ± 0.13, and showed good stability during long term cultivation. To overcome poor transformation efficiency, we developed a Cre- loxP -mediated in vivo re-circularization strategy that enabled reliable delivery of the episomal plasmid. Using this improved system, we demonstrated functional episomal expression of metabolic engineering genes and multi-gene pathways for the production of triacetic acid lactone, fatty alcohols, and limonene. Finally, we leveraged this platform to establish a redesigned CRISPR system that enables seamless genome editing in R. toruloides for the first time, while also simplifying marker recycling. Together, this work establishes a long-needed episomal plasmid platform and associated CRISPR toolkit that will accelerate metabolic engineering, synthetic biology, and fundamental studies in R. toruloides .

CRISPR-Cas9

Data for Development, Optimization, and Application of an Episomal Plasmid System for Rhodotorula toruloides

Rhodotorula toruloides is an emerging oleaginous yeast with strong potential as a microbial cell factory for the production of acetyl-CoA-derived bioproducts. However, engineering of this organism has been limited by the absence of a functional episomal plasmid system, a foundational genetic tool for rapid gene expression, pathway testing, and CRISPR-based genome engineering. Here, we report the first episomal plasmid system for R. toruloides . Through systematic screening of candidate autonomously replicating sequences (ARSs) from diverse sources, we identified multiple functional ARS elements and selected C63F4, a fragment derived from Contig 63 of R. toruloides CBS14, because of its stable performance. The resulting pC63F4 plasmid was maintained episomally, supported GFP reporter expression, exhibited a copy number of 2.39 ± 0.13, and showed good stability during long term cultivation. To overcome poor transformation efficiency, we developed a Cre-loxP-mediated in vivo re-circularization strategy that enabled reliable delivery of the episomal plasmid. Using this improved system, we demonstrated functional episomal expression of metabolic engineering genes and multi-gene pathways for the production of triacetic acid lactone, fatty alcohols, and limonene. Finally, we leveraged this platform to establish a redesigned CRISPR system that enables seamless genome editing in R. toruloides for the first time, while also simplifying marker recycling. Together, this work establishes a long-needed episomal plasmid platform and associated CRISPR toolkit that will accelerate metabolic engineering, synthetic biology, and fundamental studies in R. toruloides .

Gene Editing

Effect Of Spaceflight On Microbial Gene Expression And Virulence: Preliminary Results From Microbe Payload Flown On-Board STS-115

Human presence in space, whether permanent or temporary, is accompanied by the presence of microbes. However, the extent of microbial changes in response to spaceflight conditions and the corresponding changes to infectious disease risk is unclear. Previous studies have indicated that spaceflight weakens the immune system in humans and animals. In addition, preflight and in-flight monitoring of the International Space Station (ISS) and other spacecraft indicates the presence of opportunistic pathogens and the potential of obligate pathogens. Altered antibiotic resistance of microbes in flight has also been shown. As astronauts and cosmonauts live for longer periods in a closed environment, especially one using recycled water and air, there is an increased risk to crewmembers of infectious disease events occurring in-flight. Therefore, understanding how the space environment affects microorganisms and their disease potential is critically important for spaceflight missions and requires further study. The goal of this flight experiment, operationally called MICROBE, is to utilize three model microbial pathogens, Salmonella typhimurium, Pseudomonas aeruginosa, and Candida albicans to examine the global effects of spaceflight on microbial gene expression and virulence attributes. Specifically, the aims are (1) to perform microarray-mediated gene expression profiling of S. typhimurium, P. aeruginosa, and C. albicans, in response to spaceflight in comparison to ground controls and (2) to determine the effect of spaceflight on the virulence potential of these microorganisms immediately following their return from spaceflight using murine models. The model microorganisms were selected as they have been isolated from preflight or in-flight monitoring, represent different degrees of pathogenic behavior, are well characterized, and have sequenced genomes with available microarrays. In particular, extensive studies of S. typhimurium by the Principal Investigator, Dr. Nickerson, using ground-based analog systems demonstrate important changes in the genotypic, phenotypic, and virulence characteristics of this pathogen resulting from exposure to a flight-like environment (i.e. modeled microgravity).

Wilson, J. W.