Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Biological Data for Deep Space Mission Support

Increased biomedical risks and challenges associated with deep space missions (cis-Lunar, Mars transit, Mars surface) require new knowledge discovery and development of novel ecosystem and biomedical support capabilities. This paradigm shift supporting distant and long-duration missions requires biological data to be findable, accessible, interoperable, reusable (FAIR), and maximally open-access (i.e., there is a data governance continuum from closed to mediated to embargoed to open). The NASA “Open Science Data Repositories” (OSDR) aims to meet scientific, technical, and operational spaceflight needs, and offers the ability to upload, download, search, share, analyze, and visualize data across physiological, behavioral, ‘omics, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive (ALSDA), and NASA Biological Institutional Scientific Collection (NBISC). In the past year, ALSDA has undergone a transformation in its data collection, curation, and architecture methods. Standardizing non-genomic (phenotypic) datasets was, and will continue to be, a challenge because of their diverse nature (e.g., molecular, cellular, tissue, whole organism, behavior; micro-computed tomography, intraocular pressure, fluorescence microscopy, western blot, ultrasonography; tabular, images, video). This year ALSDA, alongside GeneLab, introduced the Biological Data Management Environment (BDME) with the purpose to accept submission of data from space relevant experiments including spaceflight, radiation, simulated gravity, gravitropism, isolation and confinement, hostile closed environments and/or distance from Earth. In addition to bringing together omics, phenotypic, physiological, bioimaging, and behavioral data into one repository. By integrating with GeneLab a multi-project submission portal aims to reduce the burden on PIs submitting data and enabling the discovery of both omics and phenotypic data. The purpose of ALSDA is to collect, curate, and make all non-human space-relevant biological data maximally findable, accessible, interoperable, and reusable (FAIR). These scope of ALSDA data collected and submitted by PIs include study design metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). In 2021, a community of researchers rallied to form the ALSDA Analysis Working Group (AWG) and provided scientific consensus on dataset sample and assay metadata. The community and excitement around the ALSDA/OSDR system has already led to several data reuse studies, demonstrating value using machine learning (ML), knowledge graphs, and meta-analysis approaches.

space biology↗

Updated Virophage Taxonomy and Distinction from Polinton-like Viruses

Virophages are small dsDNA viruses that hijack the machinery of giant viruses during the co-infection of a protist (i.e., microeukaryotic) host and represent an exceptional case of “hyperparasitism” in the viral world. While only a handful of virophages have been isolated, a vast diversity of virophage-like sequences have been uncovered from diverse metagenomes. Their wide ecological distribution, idiosyncratic infection and replication strategy, ability to integrate into protist and giant virus genomes and potential role in antiviral defense have made virophages a topic of broad interest. However, one limitation for further studies is the lack of clarity regarding the nomenclature and taxonomy of this group of viruses. Specifically, virophages have been linked in the literature to other “virophage-like” mobile genetic elements and viruses, including polinton-like viruses (PLVs), but there are no formal demarcation criteria and proper nomenclature for either group, i.e., virophage or PLVs. Here, as part of the ICTV Virophage Study Group, we leverage a large set of genomes gathered from published datasets as well as newly generated protist genomes to propose delineation criteria and classification methods at multiple taxonomic ranks for virophages ‘sensu stricto’, i.e., genomes related to the prototype isolates Sputnik and mavirus. Based on a combination of comparative genomics and phylogenetic analyses, we show that this group of virophages forms a cohesive taxon that we propose to establish at the class level and suggest a subdivision into four orders and seven families with distinctive ecogenomic features. Finally, to illustrate how the proposed delineation criteria and classification method would be used, we apply these to two recently published datasets, which we show include both virophages and other virophage-related elements. Overall, we see this proposed classification as a necessary first step to provide a robust taxonomic framework in this area of the virosphere, which will need to be expanded in the future to cover other virophage-related viruses such as PLVs.

59 BASIC BIOLOGICAL SCIENCES↗

A novel candidate hepatitis C virus genotype 4 subtype identified by next generation sequencing full-genome characterization in a patient from Saudi Arabia

Background and aim: Hepatitis C virus (HCV) infection is a major global public health concern, being a leading cause of chronic liver diseases such as chronic hepatitis, cirrhosis, and hepatocellular carcinoma. The virus is classified into 8 genotypes and 93 subtypes, each displaying distinct geographic distributions. Genotype 4 is the most predominant in the Middle East and Eastern Mediterranean and is associated with high rates of hepatitis C infection worldwide. This study used next-generation sequencing to fully characterize the HCV genome and identify a novel subtype within genotype 4 isolated from a 64-year-old Saudi man diagnosed with hepatitis C. Methods: We analyzed the complete genome of the 141-HCV isolate using whole-genome sequencing. Results: Our phylogenetic reconstructions, based on the entire genome of HCV-4 strains, revealed that the 141-HCV isolate formed a distinct group within the genotype 4 classification, providing valuable new insights into the variability of HCV. Conclusion: This discovery of a previously unclassified HCV subtype within genotype 4 sheds light on the ongoing evolution and diversity of the virus. Such knowledge has significant implications for diagnostic and therapeutic approaches, as different subtypes may exhibit varying drug sensitivities and resistance profiles.

60 APPLIED LIFE SCIENCES↗

A pipeline for targeted metagenomics of environmental bacteria

Background:Metagenomics and single cell genomics provide a window into the genetic repertoire of yet uncultivated microorganisms, but both methods are usually taxonomically untargeted. The combination of fluorescence in situ hybridization (FISH) and fluorescence activated cell sorting (FACS) has the potential to enrich taxonomically well-defined clades for genomic analyses. Methods:Cells hybridized with a taxon-specific FISH probe are enriched based on their fluorescence signal via flow cytometric cell sorting. A recently developed FISH procedure, the hybridization chain reaction (HCR)-FISH, provides the high signal intensities required for flow cytometric sorting while maintaining the integrity of the cellular DNA for subsequent genome sequencing. Sorted cells are subjected to shotgun sequencing, resulting in targeted metagenomes of low diversity. Results: Pure cultures of different taxonomic groups were used to (1) adapt and optimize the HCR-FISH protocol and (2) assess the effects of various cell fixation methods on both the signal intensity for cell sorting and the quality of subsequent genome amplification and sequencing. Best results were obtained for ethanol-fixed cells in terms of both HCR-FISH signal intensity and genome assembly quality. Our newly developed pipeline was successfully applied to a marine plankton sample from the North Sea yielding good quality metagenome assembled genomes from a yet uncultivated flavobacterial clade. Conclusions: With the developed pipeline, targeted metagenomes at various taxonomic levels can be efficiently retrieved from environmental samples. The resulting metagenome assembled genomes allow for the description of yet uncharacterized microbial clades.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting the Identities of su(met-2) and met-3 in Neurospora crassa by Genome Resequencing

A significant number of classical genetic Neurospora crassa biochemical mutants remain anonymous, unassociated with a physical genome locus. By utilizing short read next-generation sequencing methods, it is possible to sequence the genomes of mutant strains rapidly and economically for the purpose of identifying genes associated with mutant phenotypes. We have taken this approach to connect genes and mutations to “methionineless” phenotypes in N. crassa.

59 BASIC BIOLOGICAL SCIENCES↗

Selective Whole-Genome Amplification as a Tool to Enrich Specimens with Low Treponema pallidum Genomic DNA Copies for Whole-Genome Sequencing

Downstream next-generation sequencing (NGS) of the syphilis spirochete Treponema pallidum subspecies pallidum (T. pallidum) is hindered by low bacterial loads and the overwhelming presence of background metagenomic DNA in clinical specimens. In this study, we investigated selective whole-genome amplification (SWGA) utilizing multiple displacement amplification (MDA) in conjunction with custom oligonucleotides with an increased specificity for the T. pallidum genome and the capture and removal of 5'-C-phosphate-G-3' (CpG) methylated host DNA using the NEBNext Microbiome DNA enrichment kit followed by MDA with the REPLI-g single cell kit as enrichment methods to improve the yields of T. pallidum DNA in isolates and lesion specimens from syphilis patients. Sequencing was performed using the Illumina MiSeq v2 500 cycle or NovaSeq 6000 SP platform. These two enrichment methods led to 93 to 98% genome coverage at 5 reads/site in 5 clinical specimens from the United States and rabbit-propagated isolates, containing >14 T. pallidum genomic copies/μL of sample for SWGA and >129 genomic copies/μL for CpG methylation capture with MDA. Variant analysis using sequencing data derived from SWGA-enriched specimens showed that all 5 clinical strains had the A2058G mutation associated with azithromycin resistance. SWGA is a robust method that allows direct whole-genome sequencing (WGS) of specimens containing very low numbers of T. pallidum, which has been challenging until now.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of genotype‐calling methodologies on genome‐wide association and genomic prediction in polyploids

Abstract Discovery and analysis of genetic variants underlying agriculturally important traits are key to molecular breeding of crops. Reduced representation approaches have provided cost‐efficient genotyping using next‐generation sequencing. However, accurate genotype calling from next‐generation sequencing data is challenging, particularly in polyploid species due to their genome complexity. Recently developed Bayesian statistical methods implemented in available software packages, polyRAD, EBG, and updog, incorporate error rates and population parameters to accurately estimate allelic dosage across any ploidy. We used empirical and simulated data to evaluate the three Bayesian algorithms and demonstrated their impact on the power of genome‐wide association study (GWAS) analysis and the accuracy of genomic prediction. We further incorporated uncertainty in allelic dosage estimation by testing continuous genotype calls and comparing their performance to discrete genotypes in GWAS and genomic prediction. We tested the genotype‐calling methods using data from two autotetraploid species, Miscanthus sacchariflorus and Vaccinium corymbosum , and performed GWAS and genomic prediction. In the empirical study, the tested Bayesian genotype‐calling algorithms differed in their downstream effects on GWAS and genomic prediction, with some showing advantages over others. Through subsequent simulation studies, we observed that at low read depth, polyRAD was advantageous in its effect on GWAS power and limit of false positives. Additionally, we found that continuous genotypes increased the accuracy of genomic prediction, by reducing genotyping error, particularly at low sequencing depth. Our results indicate that by using the Bayesian algorithm implemented in polyRAD and continuous genotypes, we can accurately and cost‐efficiently implement GWAS and genomic prediction in polyploid crops.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmark datasets for SARS-CoV-2 surveillance bioinformatics

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the cause of coronavirus disease 2019 (COVID-19), has spread globally and is being surveilled with an international genome sequencing effort. Surveillance consists of sample acquisition, library preparation, and whole genome sequencing. This has necessitated a classification scheme detailing Variants of Concern (VOC) and Variants of Interest (VOI), and the rapid expansion of bioinformatics tools for sequence analysis. These bioinformatic tools are means for major actionable results: maintaining quality assurance and checks, defining population structure, performing genomic epidemiology, and inferring lineage to allow reliable and actionable identification and classification. Additionally, the pandemic has required public health laboratories to reach high throughput proficiency in sequencing library preparation and downstream data analysis rapidly. However, both processes can be limited by a lack of a standardized sequence dataset. We identified six SARS-CoV-2 sequence datasets from recent publications, public databases and internal resources. In addition, we created a method to mine public databases to identify representative genomes for these datasets. Using this novel method, we identified several genomes as either VOI/VOC representatives or non-VOI/VOC representatives. To describe each dataset, we utilized a previously published datasets format, which describes accession information and whole dataset information. Additionally, a script from the same publication has been enhanced to download and verify all data from this study.

60 APPLIED LIFE SCIENCES↗

High-throughput, single-microbe genomics with strain resolution, applied to a human gut microbiome

We present Microbe-seq, a high-throughput single-microbe method that yields strain-resolved genomes from complex microbial communities. We encapsulate individual microbes into droplets with microfluidics and liberate their DNA, which we amplify, tag with droplet-specific barcodes, and sequence. We use Microbe-seq to explore the human gut microbiome; we collect stool samples from a single individual, sequence over 20,000 microbes, and reconstruct nearly-complete genomes of almost 100 bacterial species, including several with multiple subspecies strains. We use these genomes to probe genomic signatures of microbial interactions: we reconstruct the horizontal gene transfer (HGT) network within the individual and observe far greater exchange within the same bacterial phylum than between different phyla. We probe bacteria-virus interactions; unexpectedly, we identify a significant in vivo association between crAssphage, an abundant bacteriophage, and a single strain of Bacteroides vulgatus. Microbe-seq contributes high-throughput culture-free capabilities to investigate genomic blueprints of complex microbial communities with single-microbe resolution.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic insights into local adaptation and migration success in reintroduced Coho Salmon of the Wenatchee River basin

ABSTRACT Objective Reintroduction of salmonids into regions where they have been extirpated is a common conservation strategy that is often implemented through natural recolonization, translocation of natural populations, or hatchery-based programs. Locally adapting to specific environmental conditions is critical for long-term population viability, particularly for species like Coho Salmon Oncorhynchus kisutch, which face diverse selective pressures during their migration. This study focused on the mid-Columbia River Coho Salmon reintroduction program managed by Yakama Nation Fisheries, which has successfully reintroduced Coho Salmon into the Wenatchee and Methow River basins, Washington. Notably, these populations have adapted to the longer migration route than those in the founding stock, with selection favoring individuals with an earlier arrival time and that can navigate a 15-km, high-gradient canyon to reach optimal spawning grounds. The objectives of this study were to investigate whether specific genomic regions are under selection for traits associated with return location and timing in Coho Salmon. Methods Low-coverage whole-genome resequencing data were used to screen for genomic regions associated with the phenotypes of interest. Results A weak polygenic signal in female Coho Salmon was found to be associated with return group, with a subset of candidate adaptive regions occurring across eight chromosomes. Conclusions These findings provide insights into the genomic mechanisms underlying local adaptation in reintroduced salmon populations and inform broodstock selection strategies aimed at promoting natural production and long-term population sustainability.

Horn, Rebekah L.↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

Recombinant And Mix-Infection Finder for SARS-CoV-2 sample

The scientific and public health communities responded to the COVID-19 pandemic with sample acquisition and genome sequencing on a scale that eclipsed all prior sequencing efforts. While this can only be characterized as a resounding success story that has cemented the use of genomics for epidemiological investigations for any future infectious disease outbreak, several retrospective studies are cataloging an array of lessons learned and issues that have yet to be addressed in order to realize the full potential of genomics as a routine biosurveillance tool. We have been both developing methods to accurately assess SARS-CoV-2 genomes from complex samples, and analyzing the large volumes of international data, both at the consensus level and the raw sequencing data. During the course of our investigations and similar to other groups, we have examined COVID-19 samples with signatures from multiple lineages of SARS-CoV-2 and will describe some of our findings during the development of a novel workflow that incorporates detection and reporting of potential co-infection within samples and also highlights any evidence of within-host recombination.

Lo, Chien-Chi↗

Prediction of condition-specific regulatory genes using machine learning

Recent advances in genomic technologies have generated data on large-scale protein–DNA interactions and open chromatin regions for many eukaryotic species. How to identify condition-specific functions of transcription factors using these data has become a major challenge in genomic research. To solve this problem, we have developed a method called ConSReg, which provides a novel approach to integrate regulatory genomic data into predictive machine learning models of key regulatory genes. Using Arabidopsis as a model system, we tested our approach to identify regulatory genes in data sets from single cell gene expression and from abiotic stress treatments. Our results showed that ConSReg accurately predicted transcription factors that regulate differentially expressed genes with an average auROC of 0.84, which is 23.5–25% better than enrichment-based approaches. To further validate the performance of ConSReg, we analyzed an independent data set related to plant nitrogen responses. ConSReg provided better rankings of the correct transcription factors in 61.7% of cases, which is three times better than other plant tools. We applied ConSReg to Arabidopsis single cell RNA-seq data, successfully identifying candidate regulatory genes that control cell wall formation. Our methods provide a new approach to define candidate regulatory genes using integrated genomic data in plants.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Metagenome-assembled genome distribution and key functionality highlight importance of aerobic metabolism in Svalbard permafrost

Permafrost underlies a large portion of the land in the Northern Hemisphere. It is proposed to be an extreme habitat and home for cold-adaptive microbial communities. Upon thaw permafrost is predicted to exacerbate increasing global temperature trend, where awakening microbes decompose millennia old carbon stocks. Yet our knowledge on composition, functional potential and variance of permafrost microbiome remains limited. In this study, we conducted a deep comparative metagenomic analysis through a 2 m permafrost core from Svalbard, Norway to determine key permafrost microbiome in this climate sensitive island ecosystem. To do so, we developed comparative metagenomics methods on metagenomic-assembled genomes (MAG). We found that community composition in Svalbard soil horizons shifted markedly with depth: the dominant phylum switched from Acidobacteria and Proteobacteria in top soils (active layer) to Actinobacteria, Bacteroidetes, Chloroflexi and Proteobacteria in permafrost layers. Key metabolic potential propagated through permafrost depths revealed aerobic respiration and soil organic matter decomposition as key metabolic traits. We also found that Svalbard MAGs were enriched in genes involved in regulation of ammonium, sulfur and phosphate. Here, we provide a new perspective on how permafrost microbiome is shaped to acquire resources in competitive and limited resource conditions of deep Svalbard soils.

59 BASIC BIOLOGICAL SCIENCES↗

ATLAS: a Snakemake workflow for assembly, annotation, and genomic binning of metagenome sequence data

Background: Metagenomics and metatranscriptomics studies provide valuable insight into the composition and function of microbial populations from diverse environments, however the data processing pipelines that rely on mapping reads to gene catalogs or genome databases for cultured strains yield results that underrepresent the genes and functional potential of uncultured microbes. Recent improvements in sequence assembly methods have eased the reliance on genome databases, thereby allowing the recovery of genomes from uncultured microbes. However, configuring these tools, linking them with advanced binning and annotation tools, and maintaining provenance of the processing continues to be challenging for researchers. Results: Here we present ATLAS, a software package for customizable data processing from raw sequence reads to functional and taxonomic annotations using state-of-the-art tools to assemble, annotate, quantify, and bin metagenome and metatranscriptome data. Genome-centric resolution and abundance estimates are provided for each sample in a dataset. ATLAS is written in Python and the workflow implemented in Snakemake; it operates in a Linux environment, and is compatible with Python 3.5+ and Anaconda 3+ versions. The source code for ATLAS is freely available, distributed under a BSD-3 license. Conclusions: ATLAS provides a user-friendly, modular and customizable Snakemake workflow for metagenome and metatranscriptome data processing; it is easily installable with conda and maintained as open-source on GitHub at https://github.com/metagenome-atlas/atlas.

59 BASIC BIOLOGICAL SCIENCES↗

Broad-spectrum CRISPR-Cas13a enables efficient phage genome editing

Abstract CRISPR-Cas13 proteins are RNA-guided RNA nucleases that defend against incoming RNA and DNA phages by binding to complementary target phage transcripts followed by general, non-specific RNA degradation. Here we analysed the defensive capabilities of LbuCas13a from Leptotrichia buccalis and found it to have robust antiviral activity unaffected by target phage gene essentiality, gene expression timing or target sequence location. Furthermore, we find LbuCas13a antiviral activity to be broadly effective against a wide range of phages by challenging LbuCas13a against nine E. coli phages from diverse phylogenetic groups. Leveraging the versatility and potency enabled by LbuCas13a targeting, we applied LbuCas13a towards broad-spectrum phage editing. Using a two-step phage-editing and enrichment method, we achieved seven markerless genome edits in three diverse phages with 100% efficiency, including edits as large as multi-gene deletions and as small as replacing a single codon. Cas13a can be applied as a generalizable tool for editing the most abundant and diverse biological entities on Earth.

59 BASIC BIOLOGICAL SCIENCES↗

Functional characterization of prokaryotic dark matter: the road so far and what lies ahead

Eight-hundred thousand to one trillion prokaryotic species may inhabit our planet. Yet, fewer than two-hundred thousand prokaryotic species have been described. This uncharted fraction of microbial diversity, and its undisclosed coding potential, is known as the “microbial dark matter” (MDM). Next-generation sequencing has allowed to collect a massive amount of genome sequence data, leading to unprecedented advances in the field of genomics. Still, harnessing new functional information from the genomes of uncultured prokaryotes is often limited by standard classification methods. These methods often rely on sequence similarity searches against reference genomes from cultured species. This hinders the discovery of unique genetic elements that are missing from the cultivated realm. It also contributes to the accumulation of prokaryotic gene products of unknown function among public sequence data repositories, highlighting the need for new approaches for sequencing data analysis and classification. Increasing evidence indicates that these proteins of unknown function might be a treasure trove of biotechnological potential. Here, we outline the challenges, opportunities, and the potential hidden within the functional dark matter (FDM) of prokaryotes. We also discuss the pitfalls surrounding molecular and computational approaches currently used to probe these uncharted waters, and discuss future opportunities for research and applications.

59 BASIC BIOLOGICAL SCIENCES↗

PARA: A New Platform for the Rapid Assembly of gRNA Arrays for Multiplexed CRISPR Technologies

Multiplexed CRISPR technologies have great potential for pathway engineering and genome editing. However, their applications are constrained by complex, laborious and time-consuming cloning steps. In this research, we developed a novel method, PARA, which allows for the one-step assembly of multiple guide RNAs (gRNAs) into a CRISPR vector with up to 18 gRNAs. Here, we demonstrate that PARA is capable of the efficient assembly of transfer RNA/Csy4/ribozyme-based gRNA arrays. To aid in this process and to streamline vector construction, we developed a user-friendly PARAweb tool for designing PCR primers and component DNA parts and simulating assembled gRNA arrays and vector sequences.

59 BASIC BIOLOGICAL SCIENCES↗