Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Accuracy-Based Annotation Quality Score (ABAQS) v1.0

Assessing genome annotation quality is crucial for downstream analyses, but current methods are inadequate for eukaryotes. We present Accuracy-Based Annotation Quality Score (ABAQS), a novel, minimal-data-driven method that comprehensively assesses annotation quality. ABAQS evaluates multiple factors, including genome completeness, gene model validity, and protein profile accuracy, outperforming other metrics like BUSCO and PSAURON. We applied ABAQS to over 2500 eukaryotic genomes and showed its robustness and effectiveness in evaluating genome annotation quality, making it a valuable tool for researchers working with genomic data. ABAQS reveals significant variation in annotation quality and highlights the importance of filtering in improving annotation quality and accuracy.

Haridas, Sajeet [Lawrence Berkeley National Labora↗

Addressing the pervasive scarcity of structural annotation in eukaryotic algae

Abstract Despite a continuous increase in algal genome sequencing, structural annotations of most algal genome assemblies remain unavailable. This pervasive scarcity of genome annotation has restricted rigorous investigation of these genomic resources and may have precipitated misleading biological interpretations. However, the annotation process for eukaryotic algal species is often challenging as genomic resources and transcriptomic evidence are not always available. To address this challenge, we benchmark the cutting-edge gene prediction methods that can be generalized for a broad range of non-model eukaryotes. Using the most accurate methods selected based on high-quality algal genomes, we predict structural annotations for 135 unannotated algal genomes. Using previously available genomic data pooled together with new data obtained in this study, we identified the core orthologous genes and the multi-gene phylogeny of eukaryotic algae, including of previously unexplored algal species. This study not only provides a benchmark for the use of structural annotation methods on a variety of non-model eukaryotes, but also compensates for missing data in the current spectrum of algal genomic resources. These results bring us one step closer to the full potential of eukaryotic algal genomics.

59 BASIC BIOLOGICAL SCIENCES↗

Biases in genome reconstruction from metagenomic data

Background Advances in sequencing, assembly, and assortment of contigs into species-specific bins has enabled the reconstruction of genomes from metagenomic data (MAGs). Though a powerful technique, it is difficult to determine whether assembly and binning techniques are accurate when applied to environmental metagenomes due to a lack of complete reference genome sequences against which to check the resulting MAGs. Methods We compared MAGs derived from an enrichment culture containing ~20 organisms to complete genome sequences of 10 organisms isolated from the enrichment culture. Factors commonly considered in binning software—nucleotide composition and sequence repetitiveness—were calculated for both the correctly binned and not-binned regions. This direct comparison revealed biases in sequence characteristics and gene content in the not-binned regions. Additionally, the composition of three public data sets representing MAGs reconstructed from the Tara Oceans metagenomic data was compared to a set of representative genomes available through NCBI RefSeq to verify that the biases identified were observable in more complex data sets and using three contemporary binning software packages. Results Repeat sequences were frequently not binned in the genome reconstruction processes, as were sequence regions with variant nucleotide composition. Genes encoded on the not-binned regions were strongly biased towards ribosomal RNAs, transfer RNAs, mobile element functions and genes of unknown function. Our results support genome reconstruction as a robust process and suggest that reconstructions determined to be >90% complete are likely to effectively represent organismal function; however, population-level genotypic heterogeneity in natural populations, such as uneven distribution of plasmids, can lead to incorrect inferences.

54 ENVIRONMENTAL SCIENCES↗

A Targeted Sequencing Assay for Serotyping Escherichia coli Using AgriSeq Technology

The gold standard method for serotyping Escherichia coli has relied on antisera-based typing of the O- and H-antigens, which is labor intensive and often unreliable. In the post-genomic era, sequence-based assays are potentially faster to provide results, could combine O-serogrouping and H-typing in a single test, and could simultaneously screen for the presence of other genetic markers of interest such as virulence factors. Whole genome sequencing is one approach; however, this method has limited multiplexing capabilities, and only a small fraction of the sequence is informative for subtyping or identifying virulence potential. A targeted, sequence-based assay and accompanying software for data analysis would be a great improvement over the currently available methods for serotyping. The purpose of this study was to develop a high-throughput, molecular method for serotyping E. coli by sequencing the genes that are required for production of O- and H-antigens, as well as to develop software for data analysis and serotype identification. To expand the utility of the assay, targets for the virulence factors, Shiga toxins ( stx 1 , and stx 2 ) and intimin ( eae ) were included. To validate the assay, genomic DNA was extracted from O-serogroup and H-type standard strains and from Shiga toxin-producing E. coli , the targeted regions were amplified, and then sequencing libraries were prepared from the amplified products followed by sequencing of the libraries on the Ion S5™ sequencer. The resulting sequence files were analyzed via the SeroType Caller™ software for identification of O-serogroup, H-type, and presence of stx 1 , stx 2 , and eae . We successfully identified 169 O-serogroups and 41 H-types. The assay also routinely detected the presence of stx 1a,c,d (3 of 3 strains), stx 2c−e,g (8 of 8 strains), stx 2f (1 strain), and eae (6 of 6 strains). Taken together, the high-throughput, sequence-based method presented here is a reliable alternative to antisera-based serotyping methods for E. coli .

Elder, Jacob R.↗

Synthetic hybrids of six yeast species

Abstract Allopolyploidy generates diversity by increasing the number of copies and sources of chromosomes. Many of the best-known evolutionary radiations, crops, and industrial organisms are ancient or recent allopolyploids. Allopolyploidy promotes differentiation and facilitates adaptation to new environments, but the tools to test its limits are lacking. Here we develop an iterative method of Hybrid Production (iHyPr) to combine the genomes of multiple budding yeast species, generating Saccharomyces allopolyploids of at least six species. When making synthetic hybrids, chromosomal instability and cell size increase dramatically as additional copies of the genome are added. The six-species hybrids initially grow slowly, but they rapidly regain fitness and adapt, even as they retain traits from multiple species. These new synthetic yeast hybrids and the iHyPr method have potential applications for the study of polyploidy, genome stability, chromosome segregation, and bioenergy.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

To have value, comparisons of high-throughput phenotyping methods need statistical tests of bias and variance

The gap between genomics and phenomics is narrowing. The rate at which it is narrowing, however, is being slowed by improper statistical comparison of methods. Quantification using Pearson’s correlation coefficient ( r ) is commonly used to assess method quality, but it is an often misleading statistic for this purpose as it is unable to provide information about the relative quality of two methods. Using r can both erroneously discount methods that are inherently more precise and validate methods that are less accurate. These errors occur because of logical flaws inherent in the use of r when comparing methods, not as a problem of limited sample size or the unavoidable possibility of a type I error. A popular alternative to using r is to measure the limits of agreement (LOA). However both r and LOA fail to identify which instrument is more or less variable than the other and can lead to incorrect conclusions about method quality. An alternative approach, comparing variances of methods, requires repeated measurements of the same subject, but avoids incorrect conclusions. Variance comparison is arguably the most important component of method validation and, thus, when repeated measurements are possible, variance comparison provides considerable value to these studies. Statistical tests to compare variances presented here are well established, easy to interpret and ubiquitously available. The widespread use of r has potentially led to numerous incorrect conclusions about method quality, hampering development, and the approach described here would be useful to advance high throughput phenotyping methods but can also extend into any branch of science. The adoption of the statistical techniques outlined in this paper will help speed the adoption of new high throughput phenotyping techniques by indicating when one should reject a new method, outright replace an old method or conditionally use a new method.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular Identification of Microbial Contaminants

Microorganisms can have significant impacts on the success of NASA’s missions, including the integrity of materials, the reliability of scientific results, and maintenance of crew health. Robust cleaning and sterilization protocols are currently in place in NASA facilities, but agency experts agree that microbial contamination is unavoidable and its impact on NASA’s missions and science must be minimized. Therefore, it is critical to understand: 1) what specific microorganisms are present, 2) how they may impact scientific objectives, and 3) how to select appropriate mitigation strategies. The Marshall Space Flight Center (MSFC) Planetary Protection (PP) microbiology lab historically relied solely upon enumeration of culturable microbial contamination associated with spacecraft materials or cleanrooms. However, this process is time consuming, many microbes cannot be cultured, and very few can be identified with any fidelity using NASA standard microbiological methods. The work described in this white paper includes the establishment of molecular identification capabilities at MSFC, including DNA isolation, amplification, purification, and Sanger sequencing. This capability will not only improve planetary protection efforts at MSFC (i.e. by identifying contaminating microorganisms in cleanrooms or on spacecraft) but also offers a service center-wide for the identification of contaminants that arise in other projects, processing locations, or during set up and roll out of spacecraft. This work also lays the foundation for higher throughput efforts to identify large populations of microbes across the lifetime of a project and serves as the starting point for future work into whole genome sequencing, non-culture based methods, or additional characterization studies. Ultimately, accurate identification informs appropriate mitigation strategies, increasing the chances of success for NASA’s missions and objectives.

C. D. Cassilly↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing↗

Methods for immortalization of epithelial cells

Methods for inducing non-clonal immortalization of normal epithelial cells by directly targeting the two main senescence barriers encountered by cultured epithelial cells. In finite lifespan pre-stasis human mammary epithelial cells (HMEC), the stress-associated stasis barrier was bypassed, and in post-stasis HMEC, the replicative senescence barrier, a consequence of critically shortened telomeres, was bypassed. Early passage non-clonal immortalized lines exhibited normal karyotypes. Methods of efficient HMEC immortalization, in the absence of “passenger” genomic errors, should facilitate examination of telomerase regulation and immortalization during human carcinoma progression, methods for screening for toxic and environmental effect on progression, and the development of therapeutics targeting the process of immortalization.

59 BASIC BIOLOGICAL SCIENCES↗

Simulation of Radiation-Induced DNA Damage With the Code RITRACKS

INTRODUCTION DNA damage is one of the most physiologically important effects of ionizing radiation. Clustered DNA damage events, like double-strand breaks (DSBs), have the most notable biological consequences. DNA damage types depend on both the track structure of the radiation and the spatial organization of the DNA. High linear energy transfer (LET) charged nuclei, found in galactic cosmic rays (GCR), are known to produce large numbers of complex DNA damage events. The human genome is packaged into chromatin, which can take on locus-dependent and cell type-dependent spatial conformations that correspond to epigenetic states, such as more open, extended structures in transcriptionally active chromatin. These epigenetic differences can affect DNA break patterns in response to ionizing radiation, potentially creating distinct DNA repair and signaling outcomes across the genome in different cells. MATERIAL AND METHODS The code RITRACKS (Relativistic Ion Tracks), which simulates stochastic radiation track structures and radiation chemistry, was used to model damage on isolated and histone-bound DNA by various types of ions and photons. The changes made to the code to perform radiation-induced DNA damage, and simulation results on single nucleosomes are given in our recent paper. In this work, the DNA building capabilities of RITRACKS have been extended to simulate more complex DNA structures build on the coarse-grain simulation framework meso-WLCsim. This code can sample generic chromatin fiber conformation ensembles based on the geometry of nucleosomes and mechanical properties of DNA. Using RITRACKS, we simulated the fragment length distributions (FLD) of irradiated DNA structures built using the chromatin conformations of WLCsim and obtained results representative of those obtained with Radiation-Induced Correlated Cleavage with sequencing (RICC-Seq) experiments [6]. We have also performed Fe ion and photon irradiations of K562, IMR90, BJ and RPE-1 cells at NSRL to experimentally validate results. Sample processing and data analysis are in progress and any available preliminary results will be discussed. DISCUSSION The recent updates in the code RITRACKS allow the calculation of several quantities such as the DNA damage yield and the FLD. This approach can be used to model epigenetic state-specific chromatin structure parameters to leverage the epigenetic state data available for many human cell types to infer relative DNA damage sensitivity among genomic loci.

I Plante↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Generating realistic building electrical load profiles through the Generative Adversarial Network (GAN)

Building electrical load profiles can improve understanding of building energy efficiency, demand flexibility, and building-grid interactions. Current approaches to generating load profiles are time-consuming and not capable of reflecting the dynamic and stochastic behaviors of real buildings; some approaches also trigger data privacy concerns. In this study, we proposed a novel approach for generating realistic electrical load profiles of buildings through the Generative Adversarial Network (GAN), a machine learning technique that is capable of revealing an unknown probability distribution purely from data. The proposed approach has three main steps: (1) normalizing the daily 24-hour load profiles, (2) clustering the daily load profiles with the k-means algorithm, and (3) using GAN to generate daily load profiles for each cluster. The approach was tested with an open-source database – the Building Data Genome Project. We validated the proposed method by comparing the mean, standard deviation, and distribution of key parameters of the generated load profiles with those of the real ones. The KL divergence of the generated and real load profiles are within 0.3 for majority of parameters and clusters. Additionally, results showed the load profiles generated by GAN can capture not only the general trend but also the random variations of the actual electrical loads in buildings. We report the proposed GAN approach can be used to generate building electrical load profiles, verify other load profile generation models, detect changes to load profiles, and more importantly, anonymize smart meter data for sharing, to support research and applications of grid-interactive efficient buildings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Biosynthesis of Haloterpenoids in Red Algae via Microbial-like Type I Terpene Synthases

Red algae or seaweeds produce highly distinctive halogenated terpenoid compounds, including the pentabromochlorinated monoterpene halomon that was once heralded as a promising anticancer agent. The first dedicated step in the biosynthesis of these natural product molecules is expected to be catalyzed by terpene synthase (TS) enzymes. Recent work has demonstrated an emerging class of type I TSs in red algal terpene biosynthesis. However, only one such enzyme from a notoriously haloterpenoid-producing red alga (Laurencia pacifica) has been functionally characterized and the product structure is not related to halogenated terpenoids. Herein, we report 10 new type I TSs from the red algae Portieria hornemannii, Plocamium pacificum, L. pacifica, and Laurencia subopposita that produce a diversity of halogenated mono- and sesquiterpenes. Here we used a combination of genome sequencing, terpenoid metabolomics, in vitro biochemistry, and bioinformatics to establish red algal TSs in all four species, including those associated with the selective production of key halogenated terpene precursors myrcene, trans-β-ocimene, and germacrene D-4-ol. These results expand on a small but growing number of characterized red algal TSs and offer insight into the biosynthesis of iconic halogenated algal compounds that are not without precedence elsewhere in biology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

scMicrobe PTA: near complete genomes from single bacterial cells

Microbial genomes produced by standard single-cell amplification methods are largely incomplete. Here, we show that primary template-directed amplification (PTA), a novel single-cell amplification technique, generated nearly complete genomes from three bacterial isolate species. Furthermore, taxonomically diverse genomes recovered from aquatic and soil microbiomes using PTA had a median completeness of 81%, whereas genomes from standard multiple displacement amplification-based approaches were usually <30% complete. PTA-derived genomes also included more associated viruses and biosynthetic gene clusters.

59 BASIC BIOLOGICAL SCIENCES↗