Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Identifying Genomic Islands with Deep Neural Networks

Background Horizontal gene transfer is the main source of adaptability for bacteria, through which genes are obtained from different sources including bacteria, archaea, viruses, and eukaryotes. This process promotes the rapid spread of genetic information across lineages, typically in the form of clusters of genes referred to as genomic islands (GIs). Different types of GIs exist, and are often classified by the content of their cargo genes or their means of integration and mobility. While various computational methods have been devised to detect different types of GIs, no single method is capable of detecting all types. Results We propose a method, which we call Shutter Island, that uses a deep learning model (Inception V3, widely used in computer vision) to detect genomic islands. The intrinsic value of deep learning methods lies in their ability to generalize. Via a technique called transfer learning, the model is pre-trained on a large generic dataset and then re-trained on images that we generate to represent genomic fragments. We demonstrate that this image-based approach generalizes better than the existing tools. Conclusions We used a deep neural network and an image-based approach to detect the most out of the correct GI predictions made by other tools, in addition to making novel GI predictions. The fact that the deep neural network was re-trained on only a limited number of GI datasets and then successfully generalized indicates that this approach could be applied to other problems in the field where data is still lacking or hard to curate.

Computer Vision↗

The first two chromosome‐scale genome assemblies of American hazelnut enable comparative genomic analysis of the genus Corylus

Summary The native, perennial shrub American hazelnut ( Corylus americana ) is cultivated in the Midwestern United States for its significant ecological benefits, as well as its high‐value nut crop. Implementation of modern breeding methods and quantitative genetic analyses of C. americana requires high‐quality reference genomes, a resource that is currently lacking. We therefore developed the first chromosome‐scale assemblies for this species using the accessions ‘Rush’ and ‘Winkler’. Genomes were assembled using HiFi PacBio reads and Arima Hi‐C data, and Oxford Nanopore reads and a high‐density genetic map were used to perform error correction. N50 scores are 31.9 Mb and 35.3 Mb, with 90.2% and 97.1% of the total genome assembled into the 11 pseudomolecules, for ‘Rush’ and ‘Winkler’, respectively. Gene prediction was performed using custom RNAseq libraries and protein homology data. ‘Rush’ has a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while ‘Winkler’ had corresponding scores of 96.9 and 96.5, indicating high‐quality assemblies. These two independent assemblies enable unbiased assessment of structural variation within C. americana , as well as patterns of syntenic relationships across the Corylus genus. Furthermore, we identified high‐density SNP marker sets from genotyping‐by‐sequencing data using 1343 C. americana , C. avellana and C. americana × C. avellana hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published Corylus genomes, were utilized to perform phylogenetic analysis of sporophytic self‐incompatibility (SSI) in hazelnut, providing evidence of unique molecular pathways governing self‐incompatibility in Corylus .

54 ENVIRONMENTAL SCIENCES↗

Integrated Phage-Host Prediction tool (iPHoP) v1.0.0

iPHoP is a bioinformatic tools that uses a set of approaches to predict the potential host of novel bacteriophages (viruses infecting bacteria) that are only known by their genome sequence, and not cultivated in the laboratory. Existing technologies typically rely on a single method, and the main advantage of iPHoP is its ability to integrate the results from multiple methods into a single prediction. This is of interest for microbial ecology researchers, as they often analyze novel bacteriophage genomes that they were able to assemble from metagenomes, but they don't know which bacteria these phages infect.

Roux, Simon↗

A new long-term sampling approach to viruses on surfaces

The importance of virus disease outbreaks and its prevention is of growing public concern but our understanding of virus transmission routes is limited by adequate sampling strategies. While conventional swabbing methods provide merely a microbial snapshot, an ideal sampling strategy would allow reliable collection of viral genomic data over longer time periods. This study has evaluated a new, paper-based sticker approach for collection of reliable viral genomic data over longer time periods up to 14 days and after implementation of different hygiene measures. In contrast to swabbing methods, which sample viral load present on a surface at a given time, the paper-based stickers are attached to the surface area of interest and collect viruses that would have otherwise been transferred onto that surface. The major advantage of one-side adhesive stickers is that they are permanently attachable to a variety of surfaces. Initial results demonstrate that stickers permit stable recovery characteristics, even at low virus titers. Stickers also allow reliable virus detection after implementation of routine hygiene measures and over longer periods up to 14 days. Overall, results for this new sticker approach for virus genomic data collection are encouraging, but further studies are required to confirm anticipated benefits over a range of virus types.

59 BASIC BIOLOGICAL SCIENCES↗

Transposon signatures of allopolyploid genome evolution

Hybridization brings together chromosome sets from two or more distinct progenitor species. Genome duplication associated with hybridization, or allopolyploidy, allows these chromosome sets to persist as distinct subgenomes during subsequent meioses. Here, we present a general method for identifying the subgenomes of a polyploid based on shared ancestry as revealed by the genomic distribution of repetitive elements that were active in the progenitors. This subgenome-enriched transposable element signal is intrinsic to the polyploid, allowing broader applicability than other approaches that depend on the availability of sequenced diploid relatives. We develop the statistical basis of the method, demonstrate its applicability in the well-studied cases of tobacco, cotton, and Brassica napus, and apply it to several cases: allotetraploid cyprinids, allohexaploid false flax, and allooctoploid strawberry. These analyses provide insight into the origins of these polyploids, revise the subgenome identities of strawberry, and provide perspective on subgenome dominance in higher polyploids.

59 BASIC BIOLOGICAL SCIENCES↗

Understanding local plant extinctions before it is too late: bridging evolutionary genomics with global ecology

Understanding evolutionary genomic and population processes within a species range is key to anticipating the extinction of plant species before it is too late. However, most models of biodiversity risk under global change do not account for the genetic variation and local adaptation of different populations. Population diversity is critical to understanding extinction because different populations may be more or less susceptible to global change and, if lost, would reduce the total diversity within a species. Here, two new modeling frameworks advance our understanding of extinction from a population and evolutionary angle: Rapid climate change-driven disruptions in population adaptation are predicted from associations between genomes and local climates. Furthermore, losses of population diversity from global land-use transformations are estimated by scaling relationships of species' genomic diversity with habitat area. Overall, these global eco-evolutionary methods advance the predictability – and possibly the preventability – of the ongoing extinction of plant species.

60 APPLIED LIFE SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Microbial production of advanced biofuels

Concerns over climate change have necessitated a rethinking of our transportation infrastructure. One possible alternative to carbon-polluting fossil fuels is biofuels produced by engineered microorganisms that use a renewable carbon source. Two biofuels, ethanol and biodiesel, have made inroads in displacing petroleum-based fuels, but their uptake has been limited by the amounts that can be used in conventional engines and by their cost. Further, advanced biofuels that mimic petroleum-based fuels are not limited by the amounts that can be used in existing transportation infrastructure but have had limited uptake due to costs. In this Review, we discuss engineering metabolic pathways to produce advanced biofuels, challenges with substrate and product toxicity with regard to host microorganisms and methods to engineer tolerance, and the use of functional genomics and machine learning approaches to produce advanced biofuels and prospects for reducing their costs.

59 BASIC BIOLOGICAL SCIENCES↗

Pseudomonas aeruginosa : One Health approach to deciphering hidden relationships in Northern Portugal

Abstract Aims Antimicrobial resistance in Pseudomonas aeruginosa represents a major global challenge in public and veterinary health, particularly from a One Health perspective. This study aimed to investigate antimicrobial resistance, the presence of virulence genes, and the genetic diversity of P. aeruginosa isolates from diverse sources. Methods and results The study utilized antimicrobial susceptibility testing, genomic analysis for resistance and virulence genes, and multilocus sequence typing to characterize a total of 737 P. aeruginosa isolates that were collected from humans, domestic animals, and aquatic environments in Northern Portugal. Antimicrobial resistance profiles were analyzed, and genomic approaches were employed to detect resistance and virulence genes. The study found a high prevalence of multidrug-resistant isolates, including high-risk clones such as ST244 and ST446, particularly in hospital sources and wastewater treatment plants. Key genes associated with resistance and virulence, including efflux pumps (e.g. MexA and MexB) and secretion systems (T3SS and T6SS), were identified. Conclusions This work highlights the intricate dynamics of multidrug-resistant P. aeruginosa across interconnected ecosystems in Northern Portugal. It underscores the importance of genomic studies in revealing the mechanisms of resistance and virulence, contributing to the broader understanding of resistance dynamics and informing future mitigation strategies.

de Sousa, Telma↗

PNNL-CompBio/bayesian-inference

Bayesian Metabolic Inference is a platform to integrate genome-scale multi-omics data into an iterative method to obtain optimized recommendations for optimal target compound production. This method uses a variational inference approach to approximate new posterior data to analyze in order to obtain useful correlations and control coefficients.

Kumar, Neeraj↗

Yeast Deletomics to Uncover Gadolinium Toxicity Targets and Resistance Mechanisms

Among the rare earth elements (REEs), a crucial group of metals for high-technologies. Gadolinium (Gd) is the only REE intentionally injected to human patients. The use of Gd-based contrasting agents for magnetic resonance imaging (MRI) is the primary route for Gd direct exposure and accumulation in humans. Consequently, aquatic environments are increasingly exposed to Gd due to its excretion through the urinary tract of patients following an MRI examination. The increasing number of reports mentioning Gd toxicity, notably originating from medical applications of Gd, necessitates an improved risk–benefit assessment of Gd utilizations. To go beyond toxicological studies, unravelling the mechanistic impact of Gd on humans and the ecosystem requires the use of genome-wide approaches. We used functional deletomics, a robust method relying on the screening of a knock-out mutant library of Saccharomyces cerevisiae exposed to toxic concentrations of Gd. The analysis of Gd-resistant and -sensitive mutants highlighted the cell wall, endosomes and the vacuolar compartment as cellular hotspots involved in the Gd response. Furthermore, we identified endocytosis and vesicular trafficking pathways (ESCRT) as well as sphingolipids homeostasis as playing pivotal roles mediating Gd toxicity. Finally, tens of yeast genes with human orthologs linked to renal dysfunction were identified as Gd-responsive. Therefore, the molecular and cellular pathways involved in Gd toxicity and detoxification uncovered in this study underline the pleotropic consequences of the increasing exposure to this strategic metal.

59 BASIC BIOLOGICAL SCIENCES↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

amsession/Kmer-based-Subgenome-Mapping

Hybridization brings together chromosome sets from two or more distinct progenitor species. Genome duplication associated with hybridization, or allopolyploidy, allows these chromosome sets to persist as distinct subgenomes during subsequent meioses. Here, we present a general method for identifying the subgenomes of a polyploid based on shared ancestry as revealed by the genomic distribution of repetitive elements that were active in the progenitors. This subgenome-enriched transposable element signal is intrinsic to the polyploid, allowing broader applicability than other approaches that depend on the availability of sequenced diploid relatives. We develop the statistical basis of the method, demonstrate its applicability in the well-studied cases of tobacco, cotton, and Brassica napus, and apply it to several cases: allotetraploid cyprinids, allohexaploid false flax, and allooctoploid strawberry. These analyses provide insight into the origins of these polyploids, revise the subgenome identities of strawberry, and provide perspective on subgenome dominance in higher polyploids.

Session, Adam↗

Identifying intragenic functional modules of genomic variations associated with cancer phenotypes by learning representation of association networks

Background Genome-wide Association Studies (GWAS) aims to uncover the link between genomic variation and phenotype. They have been actively applied in cancer biology to investigate associations between variations and cancer phenotypes, such as susceptibility to certain types of cancer and predisposed responsiveness to specific treatments. Since GWAS primarily focuses on finding associations between individual genomic variations and cancer phenotypes, there are limitations in understanding the mechanisms by which cancer phenotypes are cooperatively affected by more than one genomic variation. Results This paper proposes a network representation learning approach to learn associations among genomic variations using a prostate cancer cohort. The learned associations are encoded into representations that can be used to identify functional modules of genomic variations within genes associated with early- and late-onset prostate cancer. The proposed method was applied to a prostate cancer cohort provided by the Veterans Administration’s Million Veteran Program to identify candidates for functional modules associated with early-onset prostate cancer. The cohort included 33,159 prostate cancer patients, 3181 early-onset patients, and 29,978 late-onset patients. The reproducibility of the proposed approach clearly showed that the proposed approach can improve the model performance in terms of robustness. Conclusions To our knowledge, this is the first attempt to use a network representation learning approach to learn associations among genomic variations within genes. Associations learned in this way can lead to an understanding of the underlying mechanisms of how genomic variations cooperatively affect each cancer phenotype. This method can reveal unknown knowledge in the field of cancer biology and can be utilized to design more advanced cancer-targeted therapies.

60 APPLIED LIFE SCIENCES↗

Corrinoids as model nutrients to probe microbial interactions in a soil ecosystem

Earth’s soils are habitats for microbial communities that drive biogeochemical cycling, plant growth, and carbon storage and persistence. The thousands of microbial species living in soil form an intricate web of interactions involving the exchange of molecules produced by different microbes. Understanding in detail how these molecular exchanges occur and how they shape microbial communities may lead to new methods to improve soil health, bioremediation efforts, and better understanding of biogeochemical processes. The overall goal of this research is to gain a deeper knowledge of the microbial interactions that drive soil community structure. However, the high functional and genomic diversity in soil microbiomes has posed a challenge for current microbiology methods to achieve this goal. This research leverages a model group of key metabolites related to cobalamin (vitamin B 12 ), known as corrinoids, to investigate microbial interactions.

59 BASIC BIOLOGICAL SCIENCES↗

KBase Narrative - Mse_Genome_Assembly

Upon investigating the gut microbiome of mice for microbes related to Polycystic Ovary Syndrome (PCOS), our team found 2 metagenomes that could not be classified using alignment based methods. To further investigate microbial species from these samples, we used common MAGs workflow to create draft genomes which were then later taxonomically classified using GTDB-tk app.

Rastegar, Kiarash↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

ADEPT: a domain independent sequence alignment strategy for gpu architectures

Bioinformatic workflows frequently make use of automated genome assembly and protein clustering tools. At the core of most of these tools, a significant portion of execution time is spent in determining optimal local alignment between two sequences. This task is performed with the Smith-Waterman algorithm, which is a dynamic programming based method. With the advent of modern sequencing technologies and increasing size of both genome and protein databases, a need for faster Smith-Waterman implementations has emerged. Multiple SIMD strategies for the Smith-Waterman algorithm are available for CPUs. However, with the move of HPC facilities towards accelerator based architectures, a need for an efficient GPU accelerated strategy has emerged. Existing GPU based strategies have either been optimized for a specific type of characters (Nucleotides or Amino Acids) or for only a handful of application use-cases. In this paper, we present ADEPT, a new sequence alignment strategy for GPU architectures that is domain independent, supporting alignment of sequences from both genomes and proteins. Our proposed strategy uses GPU specific optimizations that do not rely on the nature of sequence. We demonstrate the feasibility of this strategy by implementing the Smith-Waterman algorithm and comparing it to similar CPU strategies as well as the fastest known GPU methods for each domain. ADEPT’s driver enables it to scale across multiple GPUs and allows easy integration into software pipelines which utilize large scale computational systems. We have shown that the ADEPT based Smith-Waterman algorithm demonstrates a peak performance of 360 GCUPS and 497 GCUPs for protein based and DNA based datasets respectively on a single GPU node (8 GPUs) of the Cori Supercomputer. Overall ADEPT shows 10x faster performance in a node-to-node comparison against a corresponding SIMD CPU implementation. ADEPT demonstrates a performance that is either comparable or better than existing GPU strategies. We demonstrated the efficacy of ADEPT in supporting existing bionformatics software pipelines by integrating ADEPT in MetaHipMer a high-performance denovo metagenome assembler and PASTIS a high-performance protein similarity graph construction pipeline. Our results show 10% and 30% boost of performance in MetaHipMer and PASTIS respectively.

59 BASIC BIOLOGICAL SCIENCES↗