Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Whole-Genome Sequencing and Bioinformatic Analysis of Environmental, Agricultural, and Human Campylobacter jejuni Isolates From East Tennessee

As a leading cause of bacterial-derived gastroenteritis worldwide, Campylobacter jejuni has a significant impact on human health in both the developed and developing worlds. Despite its prevalence as a human pathogen, the source of these infections remains poorly understood due to the mutation frequency of the organism and past limitations of whole genome analysis. Recent advances in both whole genome sequencing and computational methods have allowed for the high-resolution analysis of intraspecies diversity, leading multiple groups to postulate that these approaches may be used to identify the sources of Campylobacter jejuni infection. To address this hypothesis, our group conducted a regionally and temporally restricted sampling of agricultural and environmental Campylobacter sources and compared isolated C. jejuni genomes to those that caused human infections in the same region during the same time period. Through a network analysis comparing genomes from various sources, we found that human C. jejuni isolates clustered with those isolated from cattle and chickens, indicating these as potential sources of human infection in the region.

59 BASIC BIOLOGICAL SCIENCES↗

Tracking the ancestry of known and ‘ghost’ homeologous subgenomes in model grass Brachypodium polyploids

Unraveling the evolution of plant polyploids is a challenge when their diploid progenitor species are extinct or unknown or when genome sequences of known progenitors are unavailable. Existing subgenome identification methods cannot adequately infer the homeologous genomes that are present in the allopolyploids if they do not take into account the potential existence of unknown progenitors. We addressed this challenge in the widely distributed dysploid grass genus Brachypodium, which is a model genus for temperate cereals and biofuel grasses. We used a transcriptome-based phylogeny and newly designed subgenome detection algorithms coupled with a comparative chromosome barcoding analysis. Our phylogenomic subgenome detection pipeline was validated in Triticum allopolyploids, which have known progenitor genomes, and then used to infer the identities of three subgenomes derived from extant diploid species and four subgenomes derived from unknown diploid progenitors (ghost subgenomes) in six Brachypodium polyploids (B. mexicanum, B. boissieri, B. retusum, B. phoenicoides, B. rupestre and B. hybridum), of which five contain undescribed homeologous subgenomes. The existence of the seven Brachypodium progenitor genomes in the polyploids was confirmed by their karyotypic barcode profiles. Comparative phylogenomics of nuclear versus plastid trees allowed us to formulate hypothetical homoploid hybridizations and allo- and autopolyploidization scenarios that could have generated the six Brachypodium polyploids.

59 BASIC BIOLOGICAL SCIENCES↗

pnnl/pakman

PaKman: A Scalable Algorithm for Generating Genomic Contigs on Distributed Memory Machines. PaKman presents a fully distributed method that tackles assembly of large genomes through the combinationof a novel data-structure (PaK-Graph) and algorithmic strategies to simplify communication and I/O footprint during the assembly process.

Ghosh, Priyanka↗

Prokaryotic and eukaryotic cell-free systems for prototyping (CRADA Final Report)

CRADA FP00008491 between Berkeley Lab and Synvitrobio, Inc. (now Tierra Biosciences) validated the use of cell-free phenotyping to conduct functional genomics. Currently, most phenotyping work occurs using cellular fermentation methods and cellular techniques. This limits functional genomics throughput, which increasingly cannot handle the wealth of genetic information developed from next-generation sequencing technologies. De-risking a cell-free phenotyping approach has the advantage of increasing multiple-fold the throughput of genetic information that can be explored and expanding the $3B market for protein synthesis and characterization, leading to the accelerated development of human therapeutics and new biologically based materials.

59 BASIC BIOLOGICAL SCIENCES↗

CUT&RUN identifies centromeric DNA regions of Rhodotorula toruloides IFO0880

ABSTRACT Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

59 BASIC BIOLOGICAL SCIENCES↗

EEPD1 evolved a unique DNA clamping dimer protecting reversed replication forks

Exonuclease/endonuclease/phosphatase (EEP)-fold hydrolases are canonically monomeric phosphodiesterases exemplified by APE1, DNase I, and TDP2 nucleases. While EEP family domain containing protein 1 (EEPD1) acts in DNA stress responses, its proposed nuclease activities are enigmatic. Here, we integrate hybrid structural methods, evolution, biochemistry, cancer genomics, plus molecular and cell biology to define EEPD1 structure, assembly, and function at stalled DNA replication forks. Results imply EEPD1 surprisingly requires both unique EEP domain dimer and distinctive tandem Helix-hairpin-Helix [(HhH) 2 ] domains to clamp double-stranded (ds) DNA at reversed DNA replication forks for fork protection. Small-angle X-ray Scattering (SAXS), crystal, and cryo-EM structures unveil an unprecedented tryptophan handshake dimer, conserved interface di-Trp-Pro pocket, and adjustable “wrist” enabling an open-closed conformational switch. EEPD1 dimer cooperatively binds complex dsDNA replication fork intermediates but alone lacks nuclease activity due to loss of key EEP catalytic residues during Metazoan evolution and atmospheric oxygen buildup. Instead, EEPD1 prevents nucleolytic degradation of reversed replication forks by MRE11. Furthermore, cancer bioinformatics support oxidative damage-dependent EEPD1 association as a significant modulator of overall patient survival. Collective findings uncover unexpected EEP dimer and fork protection function in clamping, not cleaving, reversed replication forks for metazoan oxidative stress responses controlling genome stability and cancer outcomes.

Shen, Runze [Univ. of Texas, Houston, TX (United S↗

A symbiotic bacterium of Antarctic fish reveals environmental adaptability mechanisms and biosynthetic potential towards antibacterial and cytotoxic activities

Antarctic microbes are important agents for evolutionary adaptation and natural resource of bioactive compounds, harboring the particular metabolic pathways to biosynthesize natural products. However, not much is known on symbiotic microbiomes of fish in the Antarctic zone. In the present study, the culture method and whole-genome sequencing were performed. Natural product analyses were carried out to determine the biosynthetic potential. We report the isolation and identification of a symbiotic bacterium Serratia myotis L7-1, that is highly adaptive and resides within Antarctic fish, Trematomus bernacchii . As revealed by genomic analyses, Antarctic strain S. myotis L7-1 possesses carbohydrate-active enzymes (CAZymes), biosynthetic gene clusters (BGCs), stress response genes, antibiotic resistant genes (ARGs), and a complete type IV secretion system which could facilitate competition and colonization in the extreme Antarctic environment. The identification of microbiome gene clusters indicates the biosynthetic potential of bioactive compounds. Based on bioactivity-guided fractionation, serranticin was purified and identified as the bioactive compound, showing significant antibacterial and antitumor activity. The serranticin gene cluster was identified and located on the chrome. Furthermore, the multidrug resistance and strong bacterial antagonism contribute competitive advantages in ecological niches. Our results highlight the existence of a symbiotic bacterium in Antarctic fish largely represented by bioactive natural products and the adaptability to survive in the fish living in Antarctic oceans.

Xiao, Yu↗

Updated Virophage Taxonomy and Distinction from Polinton-like Viruses

Virophages are small dsDNA viruses that hijack the machinery of giant viruses during the co-infection of a protist (i.e., microeukaryotic) host and represent an exceptional case of “hyperparasitism” in the viral world. While only a handful of virophages have been isolated, a vast diversity of virophage-like sequences have been uncovered from diverse metagenomes. Their wide ecological distribution, idiosyncratic infection and replication strategy, ability to integrate into protist and giant virus genomes and potential role in antiviral defense have made virophages a topic of broad interest. However, one limitation for further studies is the lack of clarity regarding the nomenclature and taxonomy of this group of viruses. Specifically, virophages have been linked in the literature to other “virophage-like” mobile genetic elements and viruses, including polinton-like viruses (PLVs), but there are no formal demarcation criteria and proper nomenclature for either group, i.e., virophage or PLVs. Here, as part of the ICTV Virophage Study Group, we leverage a large set of genomes gathered from published datasets as well as newly generated protist genomes to propose delineation criteria and classification methods at multiple taxonomic ranks for virophages ‘sensu stricto’, i.e., genomes related to the prototype isolates Sputnik and mavirus. Based on a combination of comparative genomics and phylogenetic analyses, we show that this group of virophages forms a cohesive taxon that we propose to establish at the class level and suggest a subdivision into four orders and seven families with distinctive ecogenomic features. Finally, to illustrate how the proposed delineation criteria and classification method would be used, we apply these to two recently published datasets, which we show include both virophages and other virophage-related elements. Overall, we see this proposed classification as a necessary first step to provide a robust taxonomic framework in this area of the virosphere, which will need to be expanded in the future to cover other virophage-related viruses such as PLVs.

59 BASIC BIOLOGICAL SCIENCES↗

A novel candidate hepatitis C virus genotype 4 subtype identified by next generation sequencing full-genome characterization in a patient from Saudi Arabia

Background and aim: Hepatitis C virus (HCV) infection is a major global public health concern, being a leading cause of chronic liver diseases such as chronic hepatitis, cirrhosis, and hepatocellular carcinoma. The virus is classified into 8 genotypes and 93 subtypes, each displaying distinct geographic distributions. Genotype 4 is the most predominant in the Middle East and Eastern Mediterranean and is associated with high rates of hepatitis C infection worldwide. This study used next-generation sequencing to fully characterize the HCV genome and identify a novel subtype within genotype 4 isolated from a 64-year-old Saudi man diagnosed with hepatitis C. Methods: We analyzed the complete genome of the 141-HCV isolate using whole-genome sequencing. Results: Our phylogenetic reconstructions, based on the entire genome of HCV-4 strains, revealed that the 141-HCV isolate formed a distinct group within the genotype 4 classification, providing valuable new insights into the variability of HCV. Conclusion: This discovery of a previously unclassified HCV subtype within genotype 4 sheds light on the ongoing evolution and diversity of the virus. Such knowledge has significant implications for diagnostic and therapeutic approaches, as different subtypes may exhibit varying drug sensitivities and resistance profiles.

60 APPLIED LIFE SCIENCES↗

A pipeline for targeted metagenomics of environmental bacteria

Background:Metagenomics and single cell genomics provide a window into the genetic repertoire of yet uncultivated microorganisms, but both methods are usually taxonomically untargeted. The combination of fluorescence in situ hybridization (FISH) and fluorescence activated cell sorting (FACS) has the potential to enrich taxonomically well-defined clades for genomic analyses. Methods:Cells hybridized with a taxon-specific FISH probe are enriched based on their fluorescence signal via flow cytometric cell sorting. A recently developed FISH procedure, the hybridization chain reaction (HCR)-FISH, provides the high signal intensities required for flow cytometric sorting while maintaining the integrity of the cellular DNA for subsequent genome sequencing. Sorted cells are subjected to shotgun sequencing, resulting in targeted metagenomes of low diversity. Results: Pure cultures of different taxonomic groups were used to (1) adapt and optimize the HCR-FISH protocol and (2) assess the effects of various cell fixation methods on both the signal intensity for cell sorting and the quality of subsequent genome amplification and sequencing. Best results were obtained for ethanol-fixed cells in terms of both HCR-FISH signal intensity and genome assembly quality. Our newly developed pipeline was successfully applied to a marine plankton sample from the North Sea yielding good quality metagenome assembled genomes from a yet uncultivated flavobacterial clade. Conclusions: With the developed pipeline, targeted metagenomes at various taxonomic levels can be efficiently retrieved from environmental samples. The resulting metagenome assembled genomes allow for the description of yet uncharacterized microbial clades.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting the Identities of su(met-2) and met-3 in Neurospora crassa by Genome Resequencing

A significant number of classical genetic Neurospora crassa biochemical mutants remain anonymous, unassociated with a physical genome locus. By utilizing short read next-generation sequencing methods, it is possible to sequence the genomes of mutant strains rapidly and economically for the purpose of identifying genes associated with mutant phenotypes. We have taken this approach to connect genes and mutations to “methionineless” phenotypes in N. crassa.

59 BASIC BIOLOGICAL SCIENCES↗

Selective Whole-Genome Amplification as a Tool to Enrich Specimens with Low Treponema pallidum Genomic DNA Copies for Whole-Genome Sequencing

Downstream next-generation sequencing (NGS) of the syphilis spirochete Treponema pallidum subspecies pallidum (T. pallidum) is hindered by low bacterial loads and the overwhelming presence of background metagenomic DNA in clinical specimens. In this study, we investigated selective whole-genome amplification (SWGA) utilizing multiple displacement amplification (MDA) in conjunction with custom oligonucleotides with an increased specificity for the T. pallidum genome and the capture and removal of 5'-C-phosphate-G-3' (CpG) methylated host DNA using the NEBNext Microbiome DNA enrichment kit followed by MDA with the REPLI-g single cell kit as enrichment methods to improve the yields of T. pallidum DNA in isolates and lesion specimens from syphilis patients. Sequencing was performed using the Illumina MiSeq v2 500 cycle or NovaSeq 6000 SP platform. These two enrichment methods led to 93 to 98% genome coverage at 5 reads/site in 5 clinical specimens from the United States and rabbit-propagated isolates, containing >14 T. pallidum genomic copies/μL of sample for SWGA and >129 genomic copies/μL for CpG methylation capture with MDA. Variant analysis using sequencing data derived from SWGA-enriched specimens showed that all 5 clinical strains had the A2058G mutation associated with azithromycin resistance. SWGA is a robust method that allows direct whole-genome sequencing (WGS) of specimens containing very low numbers of T. pallidum, which has been challenging until now.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of genotype‐calling methodologies on genome‐wide association and genomic prediction in polyploids

Abstract Discovery and analysis of genetic variants underlying agriculturally important traits are key to molecular breeding of crops. Reduced representation approaches have provided cost‐efficient genotyping using next‐generation sequencing. However, accurate genotype calling from next‐generation sequencing data is challenging, particularly in polyploid species due to their genome complexity. Recently developed Bayesian statistical methods implemented in available software packages, polyRAD, EBG, and updog, incorporate error rates and population parameters to accurately estimate allelic dosage across any ploidy. We used empirical and simulated data to evaluate the three Bayesian algorithms and demonstrated their impact on the power of genome‐wide association study (GWAS) analysis and the accuracy of genomic prediction. We further incorporated uncertainty in allelic dosage estimation by testing continuous genotype calls and comparing their performance to discrete genotypes in GWAS and genomic prediction. We tested the genotype‐calling methods using data from two autotetraploid species, Miscanthus sacchariflorus and Vaccinium corymbosum , and performed GWAS and genomic prediction. In the empirical study, the tested Bayesian genotype‐calling algorithms differed in their downstream effects on GWAS and genomic prediction, with some showing advantages over others. Through subsequent simulation studies, we observed that at low read depth, polyRAD was advantageous in its effect on GWAS power and limit of false positives. Additionally, we found that continuous genotypes increased the accuracy of genomic prediction, by reducing genotyping error, particularly at low sequencing depth. Our results indicate that by using the Bayesian algorithm implemented in polyRAD and continuous genotypes, we can accurately and cost‐efficiently implement GWAS and genomic prediction in polyploid crops.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmark datasets for SARS-CoV-2 surveillance bioinformatics

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the cause of coronavirus disease 2019 (COVID-19), has spread globally and is being surveilled with an international genome sequencing effort. Surveillance consists of sample acquisition, library preparation, and whole genome sequencing. This has necessitated a classification scheme detailing Variants of Concern (VOC) and Variants of Interest (VOI), and the rapid expansion of bioinformatics tools for sequence analysis. These bioinformatic tools are means for major actionable results: maintaining quality assurance and checks, defining population structure, performing genomic epidemiology, and inferring lineage to allow reliable and actionable identification and classification. Additionally, the pandemic has required public health laboratories to reach high throughput proficiency in sequencing library preparation and downstream data analysis rapidly. However, both processes can be limited by a lack of a standardized sequence dataset. We identified six SARS-CoV-2 sequence datasets from recent publications, public databases and internal resources. In addition, we created a method to mine public databases to identify representative genomes for these datasets. Using this novel method, we identified several genomes as either VOI/VOC representatives or non-VOI/VOC representatives. To describe each dataset, we utilized a previously published datasets format, which describes accession information and whole dataset information. Additionally, a script from the same publication has been enhanced to download and verify all data from this study.

60 APPLIED LIFE SCIENCES↗

High-throughput, single-microbe genomics with strain resolution, applied to a human gut microbiome

We present Microbe-seq, a high-throughput single-microbe method that yields strain-resolved genomes from complex microbial communities. We encapsulate individual microbes into droplets with microfluidics and liberate their DNA, which we amplify, tag with droplet-specific barcodes, and sequence. We use Microbe-seq to explore the human gut microbiome; we collect stool samples from a single individual, sequence over 20,000 microbes, and reconstruct nearly-complete genomes of almost 100 bacterial species, including several with multiple subspecies strains. We use these genomes to probe genomic signatures of microbial interactions: we reconstruct the horizontal gene transfer (HGT) network within the individual and observe far greater exchange within the same bacterial phylum than between different phyla. We probe bacteria-virus interactions; unexpectedly, we identify a significant in vivo association between crAssphage, an abundant bacteriophage, and a single strain of Bacteroides vulgatus. Microbe-seq contributes high-throughput culture-free capabilities to investigate genomic blueprints of complex microbial communities with single-microbe resolution.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic insights into local adaptation and migration success in reintroduced Coho Salmon of the Wenatchee River basin

ABSTRACT Objective Reintroduction of salmonids into regions where they have been extirpated is a common conservation strategy that is often implemented through natural recolonization, translocation of natural populations, or hatchery-based programs. Locally adapting to specific environmental conditions is critical for long-term population viability, particularly for species like Coho Salmon Oncorhynchus kisutch, which face diverse selective pressures during their migration. This study focused on the mid-Columbia River Coho Salmon reintroduction program managed by Yakama Nation Fisheries, which has successfully reintroduced Coho Salmon into the Wenatchee and Methow River basins, Washington. Notably, these populations have adapted to the longer migration route than those in the founding stock, with selection favoring individuals with an earlier arrival time and that can navigate a 15-km, high-gradient canyon to reach optimal spawning grounds. The objectives of this study were to investigate whether specific genomic regions are under selection for traits associated with return location and timing in Coho Salmon. Methods Low-coverage whole-genome resequencing data were used to screen for genomic regions associated with the phenotypes of interest. Results A weak polygenic signal in female Coho Salmon was found to be associated with return group, with a subset of candidate adaptive regions occurring across eight chromosomes. Conclusions These findings provide insights into the genomic mechanisms underlying local adaptation in reintroduced salmon populations and inform broodstock selection strategies aimed at promoting natural production and long-term population sustainability.

Horn, Rebekah L.↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

Recombinant And Mix-Infection Finder for SARS-CoV-2 sample

The scientific and public health communities responded to the COVID-19 pandemic with sample acquisition and genome sequencing on a scale that eclipsed all prior sequencing efforts. While this can only be characterized as a resounding success story that has cemented the use of genomics for epidemiological investigations for any future infectious disease outbreak, several retrospective studies are cataloging an array of lessons learned and issues that have yet to be addressed in order to realize the full potential of genomics as a routine biosurveillance tool. We have been both developing methods to accurately assess SARS-CoV-2 genomes from complex samples, and analyzing the large volumes of international data, both at the consensus level and the raw sequencing data. During the course of our investigations and similar to other groups, we have examined COVID-19 samples with signatures from multiple lineages of SARS-CoV-2 and will describe some of our findings during the development of a novel workflow that incorporates detection and reporting of potential co-infection within samples and also highlights any evidence of within-host recombination.

Lo, Chien-Chi↗