Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

scifi-ATAC-seq: massive-scale single-cell chromatin accessibility sequencing using combinatorial fluidic indexing

Abstract Single-cell ATAC-seq has emerged as a powerful approach for revealing candidate cis-regulatory elements genome-wide at cell-type resolution. However, current single-cell methods suffer from limited throughput and high costs. Here, we present a novel technique called scifi-ATAC-seq, single-cell combinatorial fluidic indexing ATAC-sequencing, which combines a barcoded Tn5 pre-indexing step with droplet-based single-cell ATAC-seq using the 10X Genomics platform. With scifi-ATAC-seq, up to 200,000 nuclei across multiple samples can be indexed in a single emulsion reaction, representing an approximately 20-fold increase in throughput compared to the standard 10X Genomics workflow.

59 BASIC BIOLOGICAL SCIENCES↗

Sampling Microbial Dynamics in the Salish Sea Estuary: Evaluating Methods to Capture Cyanobacteria and Cyanophage

Introduction: Picocyanobacteria from the genera Prochlorococcus and Synechococcus thrive across the globe in aquatic environments, have relatively small genomes, and have growth dynamics regulated by both viral interactions and abiotic conditions, making them excellent model organisms for exploring host-pathogencoevolution. Methods: We developed and refined methods to sample and sequence cyanobacteria, cyanophages, and measured features of their abiotic environment. Results: The protocol described herein can successfully discriminate large-cell eukaryotic organisms, but size fractionation of picocyanobacteria appears to be affected by the presence of free DNA, multicellular structures, and abundant tycheposons. Our preferred final protocol from this exploratory effort included a combination of in-line and single vacuum flask filtrations, which reduced filtration processing time by over threefold in some cases compared to other tested methods, such as a fully in-line sequence or in-site filtrations. We successfully extracted an average of approximately 400–1200 ng for all filter fractions, with some variations between kits. Discussion: The protocol described herein can successfully discriminate large-cell eukaryotic organisms, but size fractionation of picocyanobacteria appears to be affected by the presence of free DNA, multicellular structures, and abundant tycheposons.

Salish Sea↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Accurate and complete genomes from metagenomes

Genomes are an integral component of the biological information about an organism; thus, the more complete the genome, the more informative it is. Historically, bacterial and archaeal genomes were reconstructed from pure (monoclonal) cultures, and the first reported sequences were manually curated to completion. However, the bottleneck imposed by the requirement for isolates precluded genomic insights for the vast majority of microbial life. Shotgun sequencing of microbial communities, referred to initially as community genomics and subsequently as genome-resolved metagenomics, can circumvent this limitation by obtaining metagenome-assembled genomes (MAGs); but gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of these genomes. Here, we discuss genome curation to improve and, in some cases, achieve complete (circularized, no gaps) MAGs (CMAGs). To date, few CMAGs have been generated, although notably some are from very complex systems such as soil and sediment. Through analysis of about 7000 published complete bacterial isolate genomes, we verify the value of cumulative GC skew in combination with other metrics to establish bacterial genome sequence accuracy. The analysis of cumulative GC skew identified potential misassemblies in some reference genomes of isolated bacteria and the repeat sequences that likely gave rise to them. We discuss methods that could be implemented in bioinformatic approaches for curation to ensure that metabolic and evolutionary analyses can be based on very high-quality genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Multiple Cases of Bacterial Sequence Erroneously Incorporated Into Publicly Available Chloroplast Genomes

Public sequencing databases are invaluable resources to biological researchers, but assessing data veracity as well as the curation and maintenance of such large collections of data can be challenging. Genomes of eukaryotic organelles, such as chloroplasts and other plastids, are particularly susceptible to assembly errors and misrepresentations in these databases due to their close evolutionary relationships with bacteria, which may co-occur within the same environment, as can be the case when sequencing plants. Here, based on sequence similarities with bacterial genomes, we identified several suspicious chloroplast assemblies present in the National Institutes of Health (NIH) Reference Sequence (RefSeq) collection. Investigations into these chloroplast assemblies reveal examples of erroneous integration of bacterial sequences into chloroplast ribosomal RNA (rRNA) loci, often within the rRNA genes, presumably due to the high similarity between plastid and bacterial rRNAs. The bacterial lineages identified within the examined chloroplasts as the most likely source of contamination are either known associates of plants, or co-occur in the same environmental niches as the examined plants. Modifications to the methods used to process untargeted ‘raw’ shotgun sequencing data from whole genome sequencing efforts, such as the identification and removal of bacterial reads prior to plastome assembly, could eliminate similar errors in the future.

59 BASIC BIOLOGICAL SCIENCES↗

A genomic analysis reveals the diversity of cellulosome displaying bacteria

Introduction Several species of cellulolytic bacteria display cellulosomes, massive multi-cellulase containing complexes that degrade lignocellulosic plant biomass (LCB). A greater understanding of cellulosome structure and enzyme content could facilitate the development of new microbial-based methods to produce renewable chemicals and materials. Methods To identify novel cellulosome-displaying microbes we searched 305,693 sequenced bacterial genomes for genes encoding cellulosome proteins; dockerin-fused glycohydrolases (DocGHs) and cohesin domain containing scaffoldins. Results and discussion This analysis identified 33 bacterial species with the genomic capacity to produce cellulosomes, including 10 species not previously reported to produce these complexes, such asAcetivibrio mesophilus. Cellulosome-producing bacteria primarily originate from theAcetivibrio, Ruminococcus, Ruminiclostridium, andClostridiumgenera. A rigorous analysis of their enzyme, scaffoldin, dockerin, and cohesin content reveals phylogenetically conserved features. Based on the presence of a high number of genes encoding both scaffoldins and dockerin-fused GHs, the cellulosomes inAcetivibrioandRuminococcusbacteria possess complex architectures that are populated with a large number of distinct LCB degrading GH enzymes. Their complex cellulosomes are distinguishable by their mechanism of attachment to the cell wall, the structures of their primary scaffoldins, and by how they are transcriptionally regulated. In contrast, bacteria in theRuminiclostridiumandClostridiumgenera produce ‘simple’ cellulosomes that are constructed from only a few types of scaffoldins that based on their distinct complement of GH enzymes are predicted to exhibit high and low cellulolytic activity, respectively. Collectively, the results of this study reveal conserved and divergent architectural features in bacterial cellulosomes that could be useful in guiding ongoing efforts to harness their cellulolytic activities for bio-based chemical and materials production.

Microbiology↗

Divergent selection and climate adaptation fuel genomic differentiation between sister species of Sphagnum (peat moss)

Abstract Background and Aims New plant species can evolve through the reinforcement of reproductive isolation via local adaptation along habitat gradients. Peat mosses (Sphagnaceae) are an emerging model system for the study of evolutionary genomics and have well-documented niche differentiation among species. Recent molecular studies have demonstrated that the globally distributed species Sphagnum magellanicum is a complex of morphologically cryptic lineages that are phylogenetically and ecologically distinct. Here, we describe the architecture of genomic differentiation between two sister species in this complex known from eastern North America: the northern S. diabolicum and the largely southern S. magniae. Methods We sampled plant populations from across a latitudinal gradient in eastern North America and performed whole genome and restriction-site associated DNA sequencing. These sequencing data were then analyzed computationally. Key Results Using sliding-window population genetic analyses we find that differentiation is concentrated within ‘islands’ of the genome spanning up to 400 kb that are characterized by elevated genetic divergence, suppressed recombination, reduced nucleotide diversity and increased rates of non-synonymous substitution. Sequence variants that are significantly associated with genetic structure and bioclimatic variables occur within genes that have functional enrichment for biological processes including abiotic stress response, photoperiodism and hormone-mediated signalling. Demographic modelling demonstrates that these two species diverged no more than 225 000 generations ago with secondary contact occurring where their ranges overlap. Conclusions We suggest that this heterogeneity of genomic differentiation is a result of linked selection and reflects the role of local adaptation to contrasting climatic zones in driving speciation. This research provides insight into the process of speciation in a group of ecologically important plants and strengthens our predictive understanding of how plant populations will respond as Earth’s climate rapidly changes.

58 GEOSCIENCES↗

Biosystems Design by Machine Learning

Biosystems such as enzymes, pathways, and whole cells have been increasingly explored for biotechnological applications. Yet, the intricate connectivity and complexity of biosystems pose a major hurdle in designing biosystems with desired features. As -omics and other high throughput technologies have been rapidly developed, the promise of applying machine learning (ML) techniques in biosystems design has started to become a reality. ML models enable the identification of patterns within complicated biological data across multiple scales of analysis and can augment biosystems design applications by predicting new candidates for optimized performance. ML is being used at every stage of biosystems design to help find non-obvious engineering solutions with fewer design iterations. In this review, we first describe commonly used models and modeling paradigms within ML. We then discuss some applications of these models that have already shown success in biotechnological applications. Moreover, we discuss successful applications at all scales of biosystems design, including nucleic acids, genetic circuits, proteins, pathways, genomes, and bioprocess. Lastly, we discuss some limitations of these methods and potential solutions as well as prospects of the combination of ML and biosystems design.

59 BASIC BIOLOGICAL SCIENCES↗

Flux REaction TArget Prioritization (Flux RETAP) v1

Metabolic engineering is evolving rapidly as a result of new advances in synthetic biology and automation, as well as the irruption of machine learning (ML). ML has been shown to provide the predictive power synthetic biology lacked and needed, and to be able to effectively guide the metabolic engineering process. However, current technical limitations prevent the independent application of ML approaches to metabolic engineering without the use of previous biological knowledge in the form of a prioritized list of desirable engineering targets. Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale metabolic models (GSMs) for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing metabolite production. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production in the literature accessible to us, 50% of targets that experimentally improved taxadiene production in E. coli and ~60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets which can also be utilized in ML pipelines.

Czajka, Jeffrey [Battelle Memorial Institute, Paci↗

Engineering organisms resistant to viruses and horizontally transferred genetic elements

Organisms resistant to horizontal gene transfer (HGT), and compositions and methods of use thereof are provided. The organisms are typically genomically recoded organisms (GRO), typically cells, having a genome in which at least one endogenous codon has been eliminated by reassignment of the codon to a synonymous or non-synonymous codon. The GRO typically include a recombinant expression construct for expression of at least one element of a ribosomal rescue pathway, typically lacking the eliminated codon. Typically, the disclosed cells are resistant to completed transfer and/or expression of a horizontally transferred genetic element (HTGE) from another organism compared a corresponding cell having a genome wherein the eliminated codon has not been eliminated. In some embodiments, the organism from which the HTGE is being transferred is a bacterium or a virus.

Isaacs, Farren J.↗

A deep learning approach to real-time HIV outbreak detection using genetic data

Pathogen genomic sequence data are increasingly made available for epidemiological monitoring. A main interest is to identify and assess the potential of infectious disease outbreaks. While popular methods to analyze sequence data often involve phylogenetic tree inference, they are vulnerable to errors from recombination and impose a high computational cost, making it difficult to obtain real-time results when the number of sequences is in or above the thousands. Here, we propose an alternative strategy to outbreak detection using genomic data based on deep learning methods developed for image classification. The key idea is to use a pairwise genetic distance matrix calculated from viral sequences as an image, and develop convolutional neutral network (CNN) models to classify areas of the images that show signatures of active outbreak, leading to identification of subsets of sequences taken from an active outbreak. We showed that our method is efficient in finding HIV-1 outbreaks with R0 ≥ 2.5, and overall a specificity exceeding 98% and sensitivity better than 92%. We validated our approach using data from HIV-1 CRF01 in Europe, containing both endemic sequences and a well-known dual outbreak in intravenous drug users. Our model accurately identified known outbreak sequences in the background of slower spreading HIV. Importantly, we detected both outbreaks early on, before they were over, implying that had this method been applied in real-time as data became available, one would have been able to intervene and possibly prevent the extent of these outbreaks. This approach is scalable to processing hundreds of thousands of sequences, making it useful for current and future real-time epidemiological investigations, including public health monitoring using large databases and especially for rapid outbreak identification.

59 BASIC BIOLOGICAL SCIENCES↗

Engineering self-deliverable ribonucleoproteins for genome editing in the brain

The delivery of CRISPR ribonucleoproteins (RNPs) for genome editing in vitro and in vivo has important advantages over other delivery methods, including reduced off-target and immunogenic effects. However, effective delivery of RNPs remains challenging in certain cell types due to low efficiency and cell toxicity. To address these issues, we engineer self-deliverable RNPs that can promote efficient cellular uptake and carry out robust genome editing without the need for helper materials or biomolecules. Screening of cell-penetrating peptides (CPPs) fused to CRISPR-Cas9 protein identifies potent constructs capable of efficient genome editing of neural progenitor cells. Further engineering of these fusion proteins establishes a C-terminal Cas9 fusion with three copies of A22p, a peptide derived from human semaphorin-3a, that exhibits substantially improved editing efficacy compared to other constructs. We find that self-deliverable Cas9 RNPs generate robust genome edits in clinically relevant genes when injected directly into the mouse striatum. Overall, self-deliverable Cas9 proteins provide a facile and effective platform for genome editing in vitro and in vivo.

59 BASIC BIOLOGICAL SCIENCES↗

Protocol for single-cell isolation and genome amplification of environmental microbial eukaryotes for genomic analysis

We describe environmental microbial eukaryotes (EMEs) sample collection, single-cell isolation, lysis, and genome amplification, followed by the rDNA amplification and OTU screening for recovery of high-quality species-specific genomes for de novo assembly. These protocols are part of our pipeline that also includes bioinformatic methods. The pipeline and its application on a wide range of phyla of different sample complexity are described in our complementary paper. In addition, this protocol describes optimized lysis, genome amplification, and OTU screening steps of the pipeline. For complete details on the use and execution of this protocol, please refer to Ciobanu et al. (2021).

59 BASIC BIOLOGICAL SCIENCES↗

Extensive Genome-Wide Phylogenetic Discordance Is Due to Incomplete Lineage Sorting and Not Ongoing Introgression in a Rapidly Radiated Bryophyte Genus

The relative importance of introgression for diversification has long been a highly disputed topic in speciation research and remains an open question despite the great attention it has received over the past decade. Gene flow leaves traces in the genome similar to those created by incomplete lineage sorting (ILS), and identification and quantification of gene flow in the presence of ILS is challenging and requires knowledge about the true phylogenetic relationship among the species. We use whole nuclear, plastid, and organellar genomes from 12 species in the rapidly radiated, ecologically diverse, actively hybridizing genus of peatmoss (Sphagnum) to reconstruct the species phylogeny and quantify introgression using a suite of phylogenomic methods. We found extensive phylogenetic discordance among nuclear and organellar phylogenies, as well as across the nuclear genome and the nodes in the species tree, best explained by extensive ILS following the rapid radiation of the genus rather than by postspeciation introgression. Our analyses support the idea of ancient introgression among the ancestral lineages followed by ILS, whereas recent gene flow among the species is highly restricted despite widespread interspecific hybridization known in the group. Our results contribute to phylogenomic understanding of how speciation proceeds in rapidly radiated, actively hybridizing species groups, and demonstrate that employing a combination of diverse phylogenomic methods can facilitate untangling complex phylogenetic patterns created by ILS and introgression.

59 BASIC BIOLOGICAL SCIENCES↗

Bioscience COVID Rapid Response Report

The COVID-19 disease outbreak and its impact on global health and economies have highlighted the national security threat posed by pathogens with pandemic potential and the need for rapid development of effective diagnostics and medical countermeasures. The Bioscience IA selected for funding rapid COVID LDRD project proposals that addressed critical R&D gaps in pandemic response that could be accomplished in 1-3 months with the requested funding. In total, the Bioscience IA funded nine rapid projects that addressed 1) rapid and accurate methods for SARS-CoV-2 RNA detection, 2) modeling tools to help prioritize populations for diagnostic testing, 3) bioinformatic tools to track SARS-CoV-2 genomic sequence changes over time, 4) molecular inhibitors of SARS-CoV-2 cellular infection, and 5) method for rapid staging of COVID19 disease to enable administration of more effective treatments. In addition, LDRD funded one larger project to be completed in FY21 that leverages Sandia capabilities to address the need for platform diagnostics and therapeutics that can be rapidly tailored against emerging pathogen targets.

59 BASIC BIOLOGICAL SCIENCES↗

Data Citation Explorer (DCE) v1.0

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data whether or not provenance was clearly indicated. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

Parker, Charles↗

Phenome‐to‐genome insights for evaluating root system architecture in field studies of maize

Abstract Understanding the genetic basis of root system architecture (RSA) in crops requires innovative approaches that enable both high‐throughput and precise phenotyping in field conditions. In this study, we evaluated multiple phenotyping and analytical frameworks for quantifying RSA in mature, field‐grown maize in three field experiments. We used forward and reverse genetic approaches to evaluate >1700 maize root crowns, including a diversity panel, a biparental mapping population, and maize mutant and wild‐type alleles at two known RSA genes,DEEPER ROOTING 1(DRO1) andRootless1(Rt1). We show the utility of increasing the dimensionality of traditional two‐dimensional (2D) techniques, referred to as the “2D multi‐view” method, to improve the capture of whole root system information for mapping genetic variation influencing RSA. Comparison of univariate and multivariate genome‐wide association study (GWAS) approaches revealed that multivariate traits were effective at dissecting complex RSA phenotypes and identifying pleiotropic quantitative trait loci (QTLs). Overall, three‐dimensional (3D) root models generated from X‐ray computed tomography and digital phenotyping captured a larger proportion of RSA trait variations compared to other methods of root phenotyping, as evidenced by both genome‐wide and single‐gene analyses. Among the individual root traits, root pulling force emerged as a highly heritable estimate of RSA that identified the largest number of shared QTLs with 3D phenotypes. Our study shows that integrating complementary phenotyping technologies helps to provide a more comprehensive understanding of the genetic architecture of RSA in field‐grown maize.

Genetics & Heredity↗

Eukaryotic genomes from a global metagenomic data set illuminate trophic modes and biogeography of ocean plankton

ABSTRACT Metagenomics is a powerful method for interpreting the ecological roles and physiological capabilities of mixed microbial communities. Yet, many tools for processing metagenomic data are neither designed to consider eukaryotes nor are they built for an increasing amount of sequence data. EukHeist is an automated pipeline to retrieve eukaryotic and prokaryotic metagenome-assembled genomes (MAGs) from large-scale metagenomic sequence data sets. We developed the EukHeist workflow to specifically process large amounts of both metagenomic and/or metatranscriptomic sequence data in an automated and reproducible fashion. Here, we applied EukHeist to the large-size fraction data (0.8–2,000 µm) from Tara Oceans to recover both eukaryotic and prokaryotic MAGs, which we refer to as TOPAZ (Tara Oceans Particle-Associated MAGs). The TOPAZ MAGs consisted of >900 environmentally relevant eukaryotic MAGs and >4,000 bacterial and archaeal MAGs. The bacterial and archaeal TOPAZ MAGs expand upon the phylogenetic diversity of likely particle- and host-associated taxa. We use these MAGs to demonstrate an approach to infer the putative trophic mode of the recovered eukaryotic MAGs. We also identify ecological cohorts of co-occurring MAGs, which are driven by specific environmental factors and putative host-microbe associations. These data together add to a number of growing resources of environmentally relevant eukaryotic genomic information. Complementary and expanded databases of MAGs, such as those provided through scalable pipelines like EukHeist, stand to advance our understanding of eukaryotic diversity through increased coverage of genomic representatives across the tree of life. IMPORTANCE Single-celled eukaryotes play ecologically significant roles in the marine environment, yet fundamental questions about their biodiversity, ecological function, and interactions remain. Environmental sequencing enables researchers to document naturally occurring protistan communities, without culturing bias, yet metagenomic and metatranscriptomic sequencing approaches cannot separate individual species from communities. To more completely capture the genomic content of mixed protistan populations, we can create bins of sequences that represent the same organism (metagenome-assembled genomes [MAGs]). We developed the EukHeist pipeline, which automates the binning of population-level eukaryotic and prokaryotic genomes from metagenomic reads. We show exciting insight into what protistan communities are present and their trophic roles in the ocean. Scalable computational tools, like EukHeist, may accelerate the identification of meaningful genetic signatures from large data sets and complement researchers’ efforts to leverage MAG databases for addressing ecological questions, resolving evolutionary relationships, and discovering potentially novel biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗