Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Draft Genome Sequences of 14 Bacterial Isolates from the Rhizosphere of Bioenergy Sorghum

We report the draft genomes of a collection of 15 bacteria, isolated from the rhizosphere soil of bioenergy sorghum (Sorghum bicolor (L.) Moench). These isolates belong to the genera, Acidovorax, Nocardioides, Agrobacterium, Peribacillus, Caulobacter, Cupriavidus, Pseudomonas, Rhizobium, Sphingomonas, Priestia, Dyadobacter, Roseomonas, Ideonella, and Bacillus.

Black, Grace S.↗

Complete genome sequence of Luteolibacter sp. strain Populi, a member of phylum Verrucomicrobiota isolated from the Populus trichocarpa rhizosphere

Luteolibacter sp. strain Populi is a bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6-Mb chromosome was completely sequenced using Oxford Nanopore long-read sequencing and is predicted to encode 5,301 proteins and 60 RNAs.

59 BASIC BIOLOGICAL SCIENCES↗

Complete genome sequence of Rhodococcus qingshengii phage Perlina

ABSTRACT Rhodococcus phage Perlina is a novel phage isolated on Rhodococcus qingshengii S10. Perlina encodes 112 open reading frames with typical phage structural genes identified and 3 tRNAs (tRNA-Ile, tRNA-Met, and tRNA-Asn). Few close relatives can be identified at the nucleotide level, suggesting a new phage species.

59 BASIC BIOLOGICAL SCIENCES↗

Develop High-Throughput Workflows for Whole-Genome Sequencing and Insertion Site Screening (CRADA Final Report)

The engineering of microbes for biomanufacturing (e.g. of fuels, chemicals, materials) applications has advanced to a stage where researchers screen genetic libraries with millions of variations each for those with enhanced productivity. This screening, however, can be slow and expensive, as screening individual variants in a high-throughput yet cost-effective manner is challenging. In this project, we aimed to reduce by 3-fold costs associated with the sequencing aspects of the screening process (to determine which genetic variant is responsible for an observed change in productivity), while being able to process over 1,000 samples per batch.

60 APPLIED LIFE SCIENCES↗

Develop High-Throughput Workflows for Whole-Genome Sequencing and Insertion Site Screening

The engineering of microbes for biomanufacturing (e.g. of fuels, chemicals, materials) applications has advanced to a stage where researchers screen genetic libraries with millions of variations each for those with enhanced productivity. This screening, however, can be slow and expensive, as screening individual variants in a high-throughput yet cost-effective manner is challenging. In this project, we aimed to reduce by 3-fold costs associated with the sequencing aspects of the screening process (to determine which genetic variant is responsible for an observed change in productivity), while being able to process over 1,000 samples per batch.

60 APPLIED LIFE SCIENCES↗

Whole-genome demography of COVID-19 virus during its pandemic period and on “panvalent” vaccine design

With over 16 million submitted genomic sequences, the SARS-CoV-2 (SC2) virus, the cause of the most recent worldwide COVID-19 pandemic, has become the most sequenced genome of all known viruses, revealing, for example, a vast number of expanding viral lineages. Since the pandemic phase appears to be over, we performed a retrospective re-examination of the demographic grouping pattern and their genomic characteristics during the entire pandemic period up to the peak of the last pandemic wave. For our study, we extracted from the NCBI only unique viral sequences and converted each sequence data to a relational vector, indicating the presence/absence of each variational event compared to a “reference” sequence. Our study revealed several genomic features that are unexpected or different from those of previous studies. For example, approximately 44,000 variants with unique sequences emerged during the pandemic period; they group into only four major viral-genomic groups and each has a set of mostly unique highly-conserved variant-genotypes (HCVGs); and a small set from the first (“ancestral”) group was inherited by the three (“descendant”) groups, suggesting that HCVGs in the next group may be predictable from the current group(s). Such a concept may be potentially important in designing “panvalent” vaccines against the current and future waves of viral infections.

60 APPLIED LIFE SCIENCES↗

Genomics and physiology of Catenibacillus, human gut bacteria capable of polyphenol C-deglycosylation and flavonoid degradation

The genusCatenibacillus(familyLachnospiraceae, phylumBacillota) includes only one cultivated species so far,Catenibacillus scindens,isolated from human faeces and capable of deglycosylating dietary polyphenols and degrading flavonoid aglycones. Another human intestinalCatenibacillusstrain not taxonomically resolved at that time was recently genome-sequenced. We analysed the genome of this novel isolate, designatedCatenibacillus decagia, and showed its ability to deglycosylateC-coupled flavone and xanthone glucosides andO-coupled flavonoid glycosides. Most of the resulting aglycones were further degraded to the corresponding phenolic acids. Including the recently sequenced genome ofC. scindensand ten faecal metagenome-assembled genomes assigned to the genusCatenibacillus, we performed a comparative genome analysis and searched for genes encoding potentialC-glycosidases and other polyphenol-converting enzymes. According to genome data and physiological characterization, the core metabolism ofCatenibacillusstrains is based on a fermentative lifestyle with butyrate production and hydrogen evolution. BothC. scindensandC. decagiaencode a flavonoidO-glycosidase, a flavone reductase, a flavanone/flavanonol-cleaving reductase and a phloretin hydrolase. Several gene clusters encode enzymes similar to those of the flavonoidC-deglycosylation system ofDoreastrain PUE (DgpBC), while separately located genes encode putative polyphenol-glucoside oxidases (DgpA) required forC-deglycosylation. The diversity ofdgpAanddgpBCgene clusters might explain the broadC-glycoside substrate spectrum ofC. scindensandC. decagia. The otherCatenibacillusgenomes encode only a few potential flavonoid-converting enzymes. Our results indicate that severalCatenibacillusspecies are well-equipped to deglycosylate and degrade dietary plant polyphenols and might inhabit a corresponding, specific niche in the gut.

Genetics & Heredity↗

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Hybridization capture sequencing for Vibrio spp. and associated virulence factors

ABSTRACT Proliferation ofVibriospp. in aquatic ecosystems is associated with climate change and, concomitantly, increased incidence of vibriosis. They are autochthonous to aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing (HCS) was employed to profile low-abundanceVibriospp. in environmental samples. The HCS panel targeted a family of molecular chaperones (CPN60) specific to 69Vibriospp. and 162Vibrio-specific virulence factors. This approach was evaluated in parallel with traditional whole-community shotgun sequencing in a metagenomic analysis of water and oyster samples collected from the Chesapeake Bay. In addition,Vibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples were subjected to whole-genome sequencing to determine the genetic characteristics of pathogenicVibriospp. circulating in an aquatic environment. HCS, employed to determine the incidence and characterization of specificVibriospp., yielded significantly greater metagenomic insight, notably a variety of otherVibriospp., including detection ofVibrio cholerae,Vibrio fluvialis, andVibrio aestuarianus, in addition toVibrio parahaemolyticusandVibrio vulnificus, and also important virulence factors not detectable using traditional molecular methods. Thus, pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood. It is concluded that environmental surveillance should include HCS, a valuable tool for the detection and characterization of pathogenic agents in aquatic ecosystems, notably vibrios. IMPORTANCE The increasing prevalence of pathogenicVibriospp. in aquatic ecosystems, driven by climate change, is closely linked to a rise in cholera and vibriosis cases, emphasizing the need for improved environmental surveillance. Vibrios are naturally occurring in aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing was employed to profile low-abundanceVibriospp. in metagenomic samples, namely water and oysters collected from the Chesapeake Bay. This approach was evaluated in parallel with traditional whole-community shotgun sequencing and whole-genome sequencing ofVibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples. Results suggest pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood, when multiple methods are considered for environmental surveillance.

Microbiology↗

Integrase-on-Demand

SAND2025-07449O Integrase-on-Demand is a software tool that allows users to identify regions in genomic sequences where genetic material can be integrated with high probability. It uses a database of integrases and their DNA attachment sites to search against any genomic sequence, producing a list of open sites, the integrase sequence, and the source of the genomic island. The program requires MASH software to be available on the system. It consists of a main script and a precomputed input file, with a taxonomy mode that searches closely related genomes and a search mode that looks for identical attachment site matches in the integrase/attachment input file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Williams, Kelly [Sandia National Lab. (SNL-CA), Li↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Genomic analysis and identification of a novel superantigen, SargEY, in Staphylococcus argenteus isolated from atopic dermatitis lesions

During surveillance of Staphylococcus aureus in lesions from patients with atopic dermatitis (AD), we isolated Staphylococcus argenteus, a species registered in 2011 as a new member of the genus Staphylococcus and previously considered a lineage of S. aureus. Genome sequence comparisons between S. argenteus isolates and representative S. aureus clinical isolates from various origins revealed that the S. argenteus genome from AD patients closely resembles that of S. aureus causing skin infections. We previously reported that 17%–22% of S. aureus isolated from skin infections produce staphylococcal enterotoxin Y (SEY), which predominantly induces T-cell proliferation via the T-cell receptor (TCR) Vα pathway. Complete genome sequencing of S. argenteus isolates revealed a gene encoding a protein similar to superantigen SEY, designated as SargEY, on its chromosome. Population structure analysis of S. argenteus revealed that these isolates are ST2250 lineage, which was the only lineage positive for the SEY-like gene among S. argenteus. Recombinant SargEY demonstrated immunological cross-reactivity with anti-SEY serum. SargEY could induce proliferation of human CD4 + and CD8 + T cells, as well as production of TNF-α and IFN-γ. SargEY showed emetic activity in a marmoset monkey model. S arg EY and SET (a phylogenetically close but uncharacterized SE) revealed their dependency on TCR Vα in inducing human T-cell proliferation. Additionally, TCR sequencing revealed other previously undescribed Vα repertoires induced by SEH. S arg EY and SEY may play roles in exacerbating the respective toxin-producing strains in AD.

59 BASIC BIOLOGICAL SCIENCES↗

Functional characterization of glycosyltransferases in duckweed to enable predictive biology

Glycosyltransferases (GTs) catalyze the formation of glycosidic linkages to produce almost all complex carbohydrates. This project used a multi-disciplinary, high-throughput (HTP) biochemical and computational biology approach focused on duckweed as a model energy crop, to study carbohydrate metabolic processes. To achieve this, developed and carried out out high-throughput (HTP) functional characterization of plant glycosyltransferases (GTs) role of enzymatic microenvironments be assessed through a combined proteomic and computational biology approach, and the combined data was used to populate deep-learning frameworks to predict plant GT function. Functional validation achieved through this research is being used to assign gene function and study plant processes at the systems level to efficiently link the genome sequence with gene function. Together, the combined approaches used within this study provide a foundation for how computational prediction, in combination with high-throughput functional validation, can be used to study plant processes at the systems level and translate knowledge gained to efficiently link genome sequence with gene function in a species agnostic manner.

09 BIOMASS FUELS↗

Revisiting synthetic lethality of Gcn5-related N-acetyltransferase (GNAT) family mutations in Haloferax volcanii

ABSTRACT Lysine acetylation is a post-translational modification that occurs in all domains of life, highlighting its evolutionary significance. Previous genome comparison identified three Gcn5-related N-acetyltransferase (GNAT) family members as lysine acetyltransferase homologs (Pat1, Pat2, and Elp3) and two deacetylase homologs (Sir2 and HdaI) in the halophilic archaeonHaloferax volcanii, withelp3andpat2proposed as a synthetic lethal gene pair. Here, we advance these findings by performing single and double mutagenesis ofelp3with thepat1andpat2lysine acetyltransferase gene homologs. Genome sequencing and PCR screens of these strains reveal successful generation of Δelp3,Δpat1Δelp3, and Δpat2Δelp3mutant strains. Although these mutant strains exhibited a reduced growth rate compared to the parent, they remained viable. Overall, this study provides genetic evidence thatelp3andpat2, while impacting cell growth, are not a synthetic lethal gene pair as previously reported. IMPORTANCE Here, we reveal by whole-genome sequencing that the GNAT family gene homologselp3andpat2can be deleted in the sameHaloferax volcaniistrain. Beyond the targeted deletions, minimal differences between the parent and Δelp3Δpat2mutant were observed, suggesting that suppressor mutations are not responsible for our ability to generate this double mutant strain. Elp3 and Pat2, thus, may not share as close a functional relationship as implied by earlier study. Our finding is significant as Elp3 is thought to function in acetylation in tRNA modification, while Pat2 likely functions in the lysine acetylation of proteins.

Microbiology↗