Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Prediction of condition-specific regulatory genes using machine learning

Recent advances in genomic technologies have generated data on large-scale protein–DNA interactions and open chromatin regions for many eukaryotic species. How to identify condition-specific functions of transcription factors using these data has become a major challenge in genomic research. To solve this problem, we have developed a method called ConSReg, which provides a novel approach to integrate regulatory genomic data into predictive machine learning models of key regulatory genes. Using Arabidopsis as a model system, we tested our approach to identify regulatory genes in data sets from single cell gene expression and from abiotic stress treatments. Our results showed that ConSReg accurately predicted transcription factors that regulate differentially expressed genes with an average auROC of 0.84, which is 23.5–25% better than enrichment-based approaches. To further validate the performance of ConSReg, we analyzed an independent data set related to plant nitrogen responses. ConSReg provided better rankings of the correct transcription factors in 61.7% of cases, which is three times better than other plant tools. We applied ConSReg to Arabidopsis single cell RNA-seq data, successfully identifying candidate regulatory genes that control cell wall formation. Our methods provide a new approach to define candidate regulatory genes using integrated genomic data in plants.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Metagenome-assembled genome distribution and key functionality highlight importance of aerobic metabolism in Svalbard permafrost

Permafrost underlies a large portion of the land in the Northern Hemisphere. It is proposed to be an extreme habitat and home for cold-adaptive microbial communities. Upon thaw permafrost is predicted to exacerbate increasing global temperature trend, where awakening microbes decompose millennia old carbon stocks. Yet our knowledge on composition, functional potential and variance of permafrost microbiome remains limited. In this study, we conducted a deep comparative metagenomic analysis through a 2 m permafrost core from Svalbard, Norway to determine key permafrost microbiome in this climate sensitive island ecosystem. To do so, we developed comparative metagenomics methods on metagenomic-assembled genomes (MAG). We found that community composition in Svalbard soil horizons shifted markedly with depth: the dominant phylum switched from Acidobacteria and Proteobacteria in top soils (active layer) to Actinobacteria, Bacteroidetes, Chloroflexi and Proteobacteria in permafrost layers. Key metabolic potential propagated through permafrost depths revealed aerobic respiration and soil organic matter decomposition as key metabolic traits. We also found that Svalbard MAGs were enriched in genes involved in regulation of ammonium, sulfur and phosphate. Here, we provide a new perspective on how permafrost microbiome is shaped to acquire resources in competitive and limited resource conditions of deep Svalbard soils.

59 BASIC BIOLOGICAL SCIENCES↗

ATLAS: a Snakemake workflow for assembly, annotation, and genomic binning of metagenome sequence data

Background: Metagenomics and metatranscriptomics studies provide valuable insight into the composition and function of microbial populations from diverse environments, however the data processing pipelines that rely on mapping reads to gene catalogs or genome databases for cultured strains yield results that underrepresent the genes and functional potential of uncultured microbes. Recent improvements in sequence assembly methods have eased the reliance on genome databases, thereby allowing the recovery of genomes from uncultured microbes. However, configuring these tools, linking them with advanced binning and annotation tools, and maintaining provenance of the processing continues to be challenging for researchers. Results: Here we present ATLAS, a software package for customizable data processing from raw sequence reads to functional and taxonomic annotations using state-of-the-art tools to assemble, annotate, quantify, and bin metagenome and metatranscriptome data. Genome-centric resolution and abundance estimates are provided for each sample in a dataset. ATLAS is written in Python and the workflow implemented in Snakemake; it operates in a Linux environment, and is compatible with Python 3.5+ and Anaconda 3+ versions. The source code for ATLAS is freely available, distributed under a BSD-3 license. Conclusions: ATLAS provides a user-friendly, modular and customizable Snakemake workflow for metagenome and metatranscriptome data processing; it is easily installable with conda and maintained as open-source on GitHub at https://github.com/metagenome-atlas/atlas.

59 BASIC BIOLOGICAL SCIENCES↗

Broad-spectrum CRISPR-Cas13a enables efficient phage genome editing

Abstract CRISPR-Cas13 proteins are RNA-guided RNA nucleases that defend against incoming RNA and DNA phages by binding to complementary target phage transcripts followed by general, non-specific RNA degradation. Here we analysed the defensive capabilities of LbuCas13a from Leptotrichia buccalis and found it to have robust antiviral activity unaffected by target phage gene essentiality, gene expression timing or target sequence location. Furthermore, we find LbuCas13a antiviral activity to be broadly effective against a wide range of phages by challenging LbuCas13a against nine E. coli phages from diverse phylogenetic groups. Leveraging the versatility and potency enabled by LbuCas13a targeting, we applied LbuCas13a towards broad-spectrum phage editing. Using a two-step phage-editing and enrichment method, we achieved seven markerless genome edits in three diverse phages with 100% efficiency, including edits as large as multi-gene deletions and as small as replacing a single codon. Cas13a can be applied as a generalizable tool for editing the most abundant and diverse biological entities on Earth.

59 BASIC BIOLOGICAL SCIENCES↗

Functional characterization of prokaryotic dark matter: the road so far and what lies ahead

Eight-hundred thousand to one trillion prokaryotic species may inhabit our planet. Yet, fewer than two-hundred thousand prokaryotic species have been described. This uncharted fraction of microbial diversity, and its undisclosed coding potential, is known as the “microbial dark matter” (MDM). Next-generation sequencing has allowed to collect a massive amount of genome sequence data, leading to unprecedented advances in the field of genomics. Still, harnessing new functional information from the genomes of uncultured prokaryotes is often limited by standard classification methods. These methods often rely on sequence similarity searches against reference genomes from cultured species. This hinders the discovery of unique genetic elements that are missing from the cultivated realm. It also contributes to the accumulation of prokaryotic gene products of unknown function among public sequence data repositories, highlighting the need for new approaches for sequencing data analysis and classification. Increasing evidence indicates that these proteins of unknown function might be a treasure trove of biotechnological potential. Here, we outline the challenges, opportunities, and the potential hidden within the functional dark matter (FDM) of prokaryotes. We also discuss the pitfalls surrounding molecular and computational approaches currently used to probe these uncharted waters, and discuss future opportunities for research and applications.

59 BASIC BIOLOGICAL SCIENCES↗

PARA: A New Platform for the Rapid Assembly of gRNA Arrays for Multiplexed CRISPR Technologies

Multiplexed CRISPR technologies have great potential for pathway engineering and genome editing. However, their applications are constrained by complex, laborious and time-consuming cloning steps. In this research, we developed a novel method, PARA, which allows for the one-step assembly of multiple guide RNAs (gRNAs) into a CRISPR vector with up to 18 gRNAs. Here, we demonstrate that PARA is capable of the efficient assembly of transfer RNA/Csy4/ribozyme-based gRNA arrays. To aid in this process and to streamline vector construction, we developed a user-friendly PARAweb tool for designing PCR primers and component DNA parts and simulating assembled gRNA arrays and vector sequences.

59 BASIC BIOLOGICAL SCIENCES↗

High-Throughput Functional Genomics for Energy Production

Functional genomics remains a foundational field for establishing genotype-phenotype relationships that enable strain engineering. High-throughput (HTP) methods accelerate the Design-Build-Test-Learn cycle that currently drives synthetic biology towards a forward engineering future. Trackable mutagenesis techniques including transposon insertion sequencing and CRISPR-Cas-mediated genome editing allow for rapid fitness profiling of a collection, or library, of mutants to discover beneficial mutations. Due to the relative speed of these experiments compared to adaptive evolution experiments, iterative rounds of mutagenesis can be implemented for next-generation metabolic engineering efforts to design complex production and tolerance phenotypes. Further, the expansion of these mutagenesis techniques to novel bacteria are opening up industrial microbes that show promise for establishing a bio-based economy.

59 BASIC BIOLOGICAL SCIENCES↗

ATCCfinder - Download and Search the ATCC Genome Portal

Much strain-specific sequence data exists in research conducted before the deployment of large sequencing repositories, making it challenging to identify and validate the identity of strains used in these studies through bioinformatics and phenotyping. The American Type Culture Collection (ATCC) is an organization that sells a wide variety of microbes with strain-level taxonomy classification and associated sequenced reference genomes. Currently, ATCC does not provide a method for searching for sequence similarity between a query sequence and their database of reference genomes. Here I propose the software ATCCfinder, which utilizes ATCC application interface software (API) to generate query-able databases from ATCC Genome resources.

Koehler, Samuel↗

Ornamental origins and genomic frontiers: a review of big-bracted dogwood research

The big-bracted (Benthamidia) dogwood clade consists of small- to medium-sized deciduous trees within the genus Cornus, known for their showy spring-time floral bract display. Cornus is within the family Cornaceae and order Cornales, and as Cornales is one of the earliest diverging asterids, these taxa have been important for phylogenetic research. Three species within the big-bracted clade, flowering (Cornus florida), kousa (C. kousa), and Pacific (C. nuttallii) dogwoods, are popular ornamental landscape plants in North America, with more than 130 cultivars released. Despite their commercial popularity, numerous research gaps have limited the expansion of fundamental research and dogwood breeding programs. In this present review, we aim to provide a thorough overview of our current understanding of 1) the phylogenetic and biogeographic context, 2) plant biology and major pests and pathogens impacting commercialization, 3) historical commercialization and propagation methods, and 4) genetic and genomic resources and how they have been implemented to understand these species. Research gaps and future directions to advance basic research and breeding of big-bracted ornamental dogwoods are discussed throughout.

Cornus florida↗

Simulating metagenomic stable isotope probing datasets with MetaSIPSim

DNA-stable isotope probing (DNA-SIP) links microorganisms to their in-situ function in diverse environmental samples. Combining DNA-SIP and metagenomics (metagenomic-SIP) allows us to link genomes from complex communities to their specific functions and improves the assembly and binning of these targeted genomes. However, empirical development of metagenomic-SIP methods is hindered by the complexity and cost of these studies. We developed a toolkit, ‘MetaSIPSim,’ to simulate sequencing read libraries for metagenomic-SIP experiments. MetaSIPSim is intended to generate datasets for method development and testing. To this end, we used MetaSIPSim generated data to demonstrate the advantages of metagenomic-SIP over a conventional shotgun metagenomic sequencing experiment. Through simulation we show that metagenomic-SIP improves the assembly and binning of isotopically labeled genomes relative to a conventional metagenomic approach. Improvements were dependent on experimental parameters and on sequencing depth. Community level G + C content impacted the assembly of labeled genomes and subsequent binning, where high community G + C generally reduced the benefits of metagenomic-SIP. Furthermore, when a high proportion of the community is isotopically labeled, the benefits of metagenomic-SIP decline. Finally, the choice of gradient fractions to sequence greatly influences method performance. Metagenomic-SIP is a valuable method for recovering isotopically labeled genomes from complex communities. We show that metagenomic-SIP performance depends on optimization of experimental parameters. MetaSIPSim allows for simulation of metagenomic-SIP datasets which facilitates the optimization and development of metagenomic-SIP experiments and analytical approaches for dealing with these data.

59 BASIC BIOLOGICAL SCIENCES↗

Deconvolute individual genomes from metagenome sequences through short read clustering

Metagenome assembly from short next-generation sequencing data is a challenging process due to its large scale and computational complexity. Clustering short reads by species before assembly offers a unique opportunity for parallel downstream assembly of genomes with individualized optimization. However, current read clustering methods suffer either false negative (under-clustering) or false positive (over-clustering) problems. Here we extended our previous read clustering software, SpaRC, by exploiting statistics derived from multiple samples in a dataset to reduce the under-clustering problem. Using synthetic and real-world datasets we demonstrated that this method has the potential to cluster almost all of the short reads from genomes with sufficient sequencing coverage. The improved read clustering in turn leads to improved downstream genome assembly quality.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Deletion of the cytochrome bc complex from Heliobacterium modesticaldum results in viable but non-phototrophic cells

The heliobacteria, a family of anoxygenic phototrophs, possess the simplest known photosynthetic apparatus. Although they are photoheterotrophs in the light, the heliobacteria can also grow chemotrophically via pyruvate metabolism in the dark. In the heliobacteria, the cytochrome bc complex is responsible for oxidizing menaquinol and reducing cytochrome c 553 in the electron flow cycle used for phototrophy. However, there is no known electron acceptor for the mobile cytochrome c 553 other than the photochemical reaction center. We have, therefore, hypothesized that the cytochrome bc complex is necessary for phototrophy, but unnecessary for chemotrophic growth in the dark. Here, we used a two-step method for CRISPR-based genome editing in Heliobacterium modesticaldum to delete the genes encoding the four major subunits of the cytochrome bc complex. Genotypic analysis verified the deletion of the petCBDA gene cluster encoding the catalytic components of the complex. Spectroscopic studies revealed that re-reduction of cytochrome c 553 after flash-induced photo-oxidation was over 100 times slower in the petCBDA mutant compared to the wild-type. Steady-state levels of oxidized P 800 (the primary donor of the photochemical reaction center) were much higher in the petCBDA mutant at every light level, consistent with a limitation in electron flow to the reaction center. The petCBDA mutant was unable to grow phototrophically on acetate plus CO 2 but could grow chemotrophically on pyruvate as a carbon source similar to the wild-type strain in the dark. The mutants could be complemented by reintroduction of the petCBDA gene cluster on a plasmid expressed from the clostridial eno promoter.

59 BASIC BIOLOGICAL SCIENCES↗

Dissecting the Shared Genetic Architecture of Suicide Attempt, Psychiatric Disorders, and Known Risk Factors

Suicide is a leading cause of death worldwide, and non-fatal suicide attempts, which occur far more frequently, are a major source of disability and social and economic burden. Both have substantial genetic etiology, which is partially shared and partially distinct from that of related psychiatric disorders. Methods: We conducted a genome-wide association study (GWAS) of 29,782 suicide attempt (SA) cases and 519,961 controls in the International Suicide Genetics Consortium. The GWAS of SA was conditioned on psychiatric disorders using GWAS summary statistics via mtCOJO, to remove genetic effects on SA mediated by psychiatric disorders. We investigated the shared and divergent genetic architectures of SA, psychiatric disorders and other known risk factors. Results: Two loci reached genome-wide significance for SA: the major histocompatibility complex and an intergenic locus on chromosome 7, which remained associated with SA after conditioning on psychiatric disorders and replicated in an independent cohort from the Million Veteran Program. This locus has been implicated in risk-taking, smoking, and insomnia. SA showed strong genetic correlation with psychiatric disorders, particularly major depression, and also with smoking, pain, risk-taking, sleep disturbances, lower educational attainment, reproductive traits, lower socioeconomic status and poorer general health. After conditioning on psychiatric disorders, the genetic correlations between SA and psychiatric disorders decreased, whereas those with non-psychiatric traits remained largely unchanged. Conclusions: Our results identify a risk locus that contributes more strongly to SA than other phenotypes and suggest a shared underlying biology between SA and known risk factors that is not mediated by psychiatric disorders.

60 APPLIED LIFE SCIENCES↗

Apomixis in Farmers’ Fields: Overview, Case Studies from Forage Grasses and Considerations for Future Apomictic Crops

Apomixis occurs naturally in several commercially important species from diverse plant families. While in some of these species apomixis is yet to be exploited in breeding schemes aimed at fixing heterosis, genetic progress and cultivar development, in other species apomixis has been integrated at different stages of breeding. Some of the most relevant examples come from the subfamily Panicoideae, the second largest subfamily of the Poaceae, and are the main focus of this review. The subfamily encompasses many tropical and sub-tropical grasses and grains of worldwide economic importance. Apomictic tropical forages are prime examples of how apomixis can be used and exploited in the development of marketable cultivars, which are essential to the meat and milk production industries globally. The main commercial forages used as grass pastures covering millions of hectares in tropical and sub-tropical regions are polyploids exhibiting gametophytic apomixis that belong to the genus Urochloa spp. (brachiariagrasses) and to the species Megathyrsus maximus (guineagrass). Buffel grass (Cenchrus ciliaris) and Paspalum spp. are other important apomictic forages bred and used in these regions. Breeding involves large germplasm collections from the centers of origin of the species, and for most of them, sexually reproducing diploid plants have been found. Chromosomically duplicated plants that maintain sexual reproduction are used in crosses with apomictic genotypes for the development and selection of cultivars to be marketed or used as progenitors in subsequent breeding cycles. The peculiarities of each genus/species breeding programs, the cultivars obtained from these programs, and the impact of use of marker assisted selection in cultivar development are presented. In addition, the test or implementation of new technologies such as high throughput phenotyping, and the use of machine learning methods for trait prediction and genomic selection are positively impacting the selection and speed of development of new polyploid apomictic cultivars. Furthermore, genetic transformation techniques, including genome editing, provide an additional layer for design of tailor-made, customer-oriented cultivars.

Cenchrus↗

Efficient DNA sequence compression with neural networks

Abstract Background The increasing production of genomic data has led to an intensified need for models that can cope efficiently with the lossless compression of DNA sequences. Important applications include long-term storage and compression-based data analysis. In the literature, only a few recent articles propose the use of neural networks for DNA sequence compression. However, they fall short when compared with specific DNA compression tools, such as GeCo2. This limitation is due to the absence of models specifically designed for DNA sequences. In this work, we combine the power of neural networks with specific DNA models. For this purpose, we created GeCo3, a new genomic sequence compressor that uses neural networks for mixing multiple context and substitution-tolerant context models. Findings We benchmark GeCo3 as a reference-free DNA compressor in 5 datasets, including a balanced and comprehensive dataset of DNA sequences, the Y-chromosome and human mitogenome, 2 compilations of archaeal and virus genomes, 4 whole genomes, and 2 collections of FASTQ data of a human virome and ancient DNA. GeCo3 achieves a solid improvement in compression over the previous version (GeCo2) of $2.4\%$, $7.1\%$, $6.1\%$, $5.8\%$, and $6.0\%$, respectively. To test its performance as a reference-based DNA compressor, we benchmark GeCo3 in 4 datasets constituted by the pairwise compression of the chromosomes of the genomes of several primates. GeCo3 improves the compression in $12.4\%$, $11.7\%$, $10.8\%$, and $10.1\%$ over the state of the art. The cost of this compression improvement is some additional computational time (1.7–3 times slower than GeCo2). The RAM use is constant, and the tool scales efficiently, independently of the sequence size. Overall, these values outperform the state of the art. Conclusions GeCo3 is a genomic sequence compressor with a neural network mixing approach that provides additional gains over top specific genomic compressors. The proposed mixing method is portable, requiring only the probabilities of the models as inputs, providing easy adaptation to other data compressors or compression-based data analysis tools. GeCo3 is released under GPLv3 and is available for free download at https://github.com/cobilab/geco3.

Silva, Milton↗

Quantification of Cas9 binding and cleavage across diverse guide sequences maps landscapes of target engagement

The RNA-guided nuclease Cas9 has unlocked powerful methods for perturbing both the genome through targeted DNA cleavage and the regulome through targeted DNA binding, but limited biochemical data have hampered efforts to quantitatively model sequence perturbation of target binding and cleavage across diverse guide sequences. We present scalable, sequencing-based platforms for high-throughput filter binding and cleavage and then perform 62,444 quantitative binding and cleavage assays on 35,047 on- and off-target DNA sequences across 90 Cas9 ribonucleoproteins (RNPs) loaded with distinct guide RNAs. We observe that binding and cleavage efficacy, as well as specificity, vary substantially across RNPs; canonically studied guides often have atypically high specificity; sequence context surrounding the target modulates Cas9 on-rate; and Cas9 RNPs may sequester targets in nonproductive states that contribute to “proofreading” capability. Lastly, we distill our findings into an interpretable biophysical model that predicts changes in binding and cleavage for diverse target sequence perturbations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Simple, Cost-Effective, and Automation-Friendly Direct PCR Approach for Bacterial Community Analysis

Understanding bacterial interactions and assembly in complex microbial communities using 16S rRNA sequencing normally requires a large experimental load. However, the current DNA extraction methods, including cell disruption and genomic DNA purification, are normally biased, costly, time-consuming, labor-intensive, and not amenable to miniaturization by droplets or 1,536-well plates due to the significant DNA loss during the purification step for tiny-volume and low-cell-density samples.

16S rRNA sequencing↗