Engineering PapersSearch

SEARCH · Engineering Papers

Results for “genome sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Whole-genome demography of COVID-19 virus during its pandemic period and on “panvalent” vaccine design

With over 16 million submitted genomic sequences, the SARS-CoV-2 (SC2) virus, the cause of the most recent worldwide COVID-19 pandemic, has become the most sequenced genome of all known viruses, revealing, for example, a vast number of expanding viral lineages. Since the pandemic phase appears to be over, we performed a retrospective re-examination of the demographic grouping pattern and their genomic characteristics during the entire pandemic period up to the peak of the last pandemic wave. For our study, we extracted from the NCBI only unique viral sequences and converted each sequence data to a relational vector, indicating the presence/absence of each variational event compared to a “reference” sequence. Our study revealed several genomic features that are unexpected or different from those of previous studies. For example, approximately 44,000 variants with unique sequences emerged during the pandemic period; they group into only four major viral-genomic groups and each has a set of mostly unique highly-conserved variant-genotypes (HCVGs); and a small set from the first (“ancestral”) group was inherited by the three (“descendant”) groups, suggesting that HCVGs in the next group may be predictable from the current group(s). Such a concept may be potentially important in designing “panvalent” vaccines against the current and future waves of viral infections.

60 APPLIED LIFE SCIENCES

Genomics and physiology of Catenibacillus, human gut bacteria capable of polyphenol C-deglycosylation and flavonoid degradation

The genusCatenibacillus(familyLachnospiraceae, phylumBacillota) includes only one cultivated species so far,Catenibacillus scindens,isolated from human faeces and capable of deglycosylating dietary polyphenols and degrading flavonoid aglycones. Another human intestinalCatenibacillusstrain not taxonomically resolved at that time was recently genome-sequenced. We analysed the genome of this novel isolate, designatedCatenibacillus decagia, and showed its ability to deglycosylateC-coupled flavone and xanthone glucosides andO-coupled flavonoid glycosides. Most of the resulting aglycones were further degraded to the corresponding phenolic acids. Including the recently sequenced genome ofC. scindensand ten faecal metagenome-assembled genomes assigned to the genusCatenibacillus, we performed a comparative genome analysis and searched for genes encoding potentialC-glycosidases and other polyphenol-converting enzymes. According to genome data and physiological characterization, the core metabolism ofCatenibacillusstrains is based on a fermentative lifestyle with butyrate production and hydrogen evolution. BothC. scindensandC. decagiaencode a flavonoidO-glycosidase, a flavone reductase, a flavanone/flavanonol-cleaving reductase and a phloretin hydrolase. Several gene clusters encode enzymes similar to those of the flavonoidC-deglycosylation system ofDoreastrain PUE (DgpBC), while separately located genes encode putative polyphenol-glucoside oxidases (DgpA) required forC-deglycosylation. The diversity ofdgpAanddgpBCgene clusters might explain the broadC-glycoside substrate spectrum ofC. scindensandC. decagia. The otherCatenibacillusgenomes encode only a few potential flavonoid-converting enzymes. Our results indicate that severalCatenibacillusspecies are well-equipped to deglycosylate and degrade dietary plant polyphenols and might inhabit a corresponding, specific niche in the gut.

Genetics & Heredity

Molecular motors and their functions in plants

Molecular motors that hydrolyze ATP and use the derived energy to generate force are involved in a variety of diverse cellular functions. Genetic, biochemical, and cellular localization data have implicated motors in a variety of functions such as vesicle and organelle transport, cytoskeleton dynamics, morphogenesis, polarized growth, cell movements, spindle formation, chromosome movement, nuclear fusion, and signal transduction. In non-plant systems three families of molecular motors (kinesins, dyneins, and myosins) have been well characterized. These motors use microtubules (in the case of kinesines and dyneins) or actin filaments (in the case of myosins) as tracks to transport cargo materials intracellularly. During the last decade tremendous progress has been made in understanding the structure and function of various motors in animals. These studies are yielding interesting insights into the functions of molecular motors and the origin of different families of motors. Furthermore, the paradigm that motors bind cargo and move along cytoskeletal tracks does not explain the functions of some of the motors. Relatively little is known about the molecular motors and their roles in plants. In recent years, by using biochemical, cell biological, molecular, and genetic approaches a few molecular motors have been isolated and characterized from plants. These studies indicate that some of the motors in plants have novel features and regulatory mechanisms. The role of molecular motors in plant cell division, cell expansion, cytoplasmic streaming, cell-to-cell communication, membrane trafficking, and morphogenesis is beginning to be understood. Analyses of the Arabidopsis genome sequence database (51% of genome) with conserved motor domains of kinesin and myosin families indicates the presence of a large number (about 40) of molecular motors and the functions of many of these motors remain to be discovered. It is likely that many more motors with novel regulatory mechanisms that perform plant-specific functions are yet to be discovered. Although the identification of motors in plants, especially in Arabidopsis, is progressing at a rapid pace because of the ongoing plant genome sequencing projects, only a few plant motors have been characterized in any detail. Elucidation of function and regulation of this multitude of motors in a given species is going to be a challenging and exciting area of research in plant cell biology. Structural features of some plant motors suggest calcium, through calmodulin, is likely to play a key role in regulating the function of both microtubule- and actin-based motors in plants.

Non-NASA Center

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES

Hybridization capture sequencing for Vibrio spp. and associated virulence factors

ABSTRACT Proliferation ofVibriospp. in aquatic ecosystems is associated with climate change and, concomitantly, increased incidence of vibriosis. They are autochthonous to aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing (HCS) was employed to profile low-abundanceVibriospp. in environmental samples. The HCS panel targeted a family of molecular chaperones (CPN60) specific to 69Vibriospp. and 162Vibrio-specific virulence factors. This approach was evaluated in parallel with traditional whole-community shotgun sequencing in a metagenomic analysis of water and oyster samples collected from the Chesapeake Bay. In addition,Vibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples were subjected to whole-genome sequencing to determine the genetic characteristics of pathogenicVibriospp. circulating in an aquatic environment. HCS, employed to determine the incidence and characterization of specificVibriospp., yielded significantly greater metagenomic insight, notably a variety of otherVibriospp., including detection ofVibrio cholerae,Vibrio fluvialis, andVibrio aestuarianus, in addition toVibrio parahaemolyticusandVibrio vulnificus, and also important virulence factors not detectable using traditional molecular methods. Thus, pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood. It is concluded that environmental surveillance should include HCS, a valuable tool for the detection and characterization of pathogenic agents in aquatic ecosystems, notably vibrios. IMPORTANCE The increasing prevalence of pathogenicVibriospp. in aquatic ecosystems, driven by climate change, is closely linked to a rise in cholera and vibriosis cases, emphasizing the need for improved environmental surveillance. Vibrios are naturally occurring in aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing was employed to profile low-abundanceVibriospp. in metagenomic samples, namely water and oysters collected from the Chesapeake Bay. This approach was evaluated in parallel with traditional whole-community shotgun sequencing and whole-genome sequencing ofVibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples. Results suggest pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood, when multiple methods are considered for environmental surveillance.

Microbiology

Integrase-on-Demand

SAND2025-07449O Integrase-on-Demand is a software tool that allows users to identify regions in genomic sequences where genetic material can be integrated with high probability. It uses a database of integrases and their DNA attachment sites to search against any genomic sequence, producing a list of open sites, the integrase sequence, and the source of the genomic island. The program requires MASH software to be available on the system. It consists of a main script and a precomputed input file, with a taxonomy mode that searches closely related genomes and a search mode that looks for identical attachment site matches in the integrase/attachment input file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Williams, Kelly [Sandia National Lab. (SNL-CA), Li

Permanent Draft Genome of Strain ESFC-1: Ecological Genomics of a Newly Discovered Lineage of Filamentous Diazotrophic Cyanobacteria

The nonheterocystous filamentous cyanobacterium, strain ESFC-1, is a recently described member of the order Oscillatoriales within the Cyanobacteria. ESFC-1 has been shown to be a major diazotroph in the intertidal microbial mat system at Elkhorn Slough, CA, USA. Based on phylogenetic analyses of the 16S RNA gene, ESFC-1 appears to belong to a unique, genus-level divergence; the draft genome sequence of this strain has now been determined. Here we report features of this genome as they relate to the ecological functions and capabilities of strain ESFC-1. The 5,632,035 bp genome sequence encodes 4914 protein-coding genes and 92 RNA genes. One striking feature of this cyanobacterium is the apparent lack of either uptake or bi-directional hydrogenases typically expected within a diazotroph. Additionally, a large genomic island is found that contains numerous low GC-content genes and genes related to extracellular polysaccharide production and cell wall synthesis and maintenance.

R Craig Everroad

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES

Discovery of Chlorophyll d : Isolation and Characterization of a Far-Red Cyanobacterium from the Original Site of Manning and Strain (1943) at Moss Beach, California

We have isolated a chlorophyll-d-containing cyanobacterium from the intertidal field site at Moss Beach, on the coast of Central California, USA, where Manning and Strain (1943) originally discovered this far-red chlorophyll. Here, we present the cyanobacterium’s environmental description, culturing procedure, pigment composition, ultrastructure, and full genome sequence. Among cultures of far-red cyanobacteria obtained from red algae from the same site, this strain was an epiphyte on a brown macroalgae. Its Q y in vivo absorbance peak is centered at 704–705 nm, the shortest wavelength observed thus far among the various known Acaryochloris strains. Its Chl a /Chl d ratio was 0.01, with Chl d accounting for 99% of the total Chl d and Chl a mass. TEM imagery indicates the absence of phycobilisomes, corroborated by both pigment spectra and genome analysis. The Moss Beach strain codes for only a single set of genes for producing allophycocyanin. Genomic sequencing yielded a 7.25 Mbp circular chromosome and 10 circular plasmids ranging from 16 kbp to 394 kbp. We have determined that this strain shares high similarity with strain S15, an epiphyte of red algae, while its distinct gene complement and ecological niche suggest that this strain could be the closest known relative to the original Chl d source of Manning and Strain (1943). The Moss Beach strain is designated Acaryochloris sp. (marina) strain Moss Beach.

chlorophyll d

Genomic analysis and identification of a novel superantigen, SargEY, in Staphylococcus argenteus isolated from atopic dermatitis lesions

During surveillance of Staphylococcus aureus in lesions from patients with atopic dermatitis (AD), we isolated Staphylococcus argenteus, a species registered in 2011 as a new member of the genus Staphylococcus and previously considered a lineage of S. aureus. Genome sequence comparisons between S. argenteus isolates and representative S. aureus clinical isolates from various origins revealed that the S. argenteus genome from AD patients closely resembles that of S. aureus causing skin infections. We previously reported that 17%–22% of S. aureus isolated from skin infections produce staphylococcal enterotoxin Y (SEY), which predominantly induces T-cell proliferation via the T-cell receptor (TCR) Vα pathway. Complete genome sequencing of S. argenteus isolates revealed a gene encoding a protein similar to superantigen SEY, designated as SargEY, on its chromosome. Population structure analysis of S. argenteus revealed that these isolates are ST2250 lineage, which was the only lineage positive for the SEY-like gene among S. argenteus. Recombinant SargEY demonstrated immunological cross-reactivity with anti-SEY serum. SargEY could induce proliferation of human CD4 + and CD8 + T cells, as well as production of TNF-α and IFN-γ. SargEY showed emetic activity in a marmoset monkey model. S arg EY and SET (a phylogenetically close but uncharacterized SE) revealed their dependency on TCR Vα in inducing human T-cell proliferation. Additionally, TCR sequencing revealed other previously undescribed Vα repertoires induced by SEH. S arg EY and SEY may play roles in exacerbating the respective toxin-producing strains in AD.

59 BASIC BIOLOGICAL SCIENCES

Gene Fusion: A Genome Wide Survey

As a well known fact, organisms form larger and complex multimodular (composite or chimeric) and mostly multi-functional proteins through gene fusion of two or more individual genes which have independent evolution histories and functions. We call each of these components a module. The existence of multimodular proteins may improves the efficiency in gene regulation and in cellular functions, and thus may give the host organism advantages in adaptation to environments. Analysis of all gene fusions in present-day organisms should allow us to examine the patterns of gene fusion in context with cellular functions, to trace back the evolution processes from the ancient smaller and uni-functional proteins to the present-day larger and complex multi-functional proteins, and to estimate the minimal number of ancestor proteins that existed in the last common ancestor for all life on earth. Although many multimodular proteins have been experimentally known, identification of gene fusion events systematically at genome scale had not been possible until recently when large number of completed genome sequences have been becoming available. In addition, technical difficulties for such analysis also exist due to the complexity of this biological and evolutionary process. We report from this study a new strategy to computationally identify multimodular proteins using completed genome sequences and the results surveyed from 22 organisms with the data from over 40 organisms to be presented during the meeting. Additional information is contained in the original extended abstract.

Liang, Ping

Spaceflight Autonomous Multigenerational Microbial Sequencer (SAMMS) in Support of Plant-Growth Systems

As the National Aeronautics and Space Association (NASA) begins to pursue long-duration space flights, they will need to be able to provide astronauts with a nutritious and reliable food source. To meet the administration’s goal of traveling to the Moon and Mars, astronauts will need to begin to grow their own food in space. To protect their food source, extensive monitoring will occur to test for the effects of a space flight environment (e.g., radiation) as well as for early pathogen and disease detection. Genomic sequencing allows for both concerns to be tested on a regular basis. However, NASA’s current sequencer is unable to process plant tissues. Therefore, a novel method for plant DNA extraction using microneedle (MN) patches that will be able to feed into NASA’s existing system, but also require minimal human input is proposed. To support this, the design was broken down into four components (1) MN patch fabrication (2) MN patch extraction, (3) automated sampling motion control, and (4) a processing module. The MN patch is fabricated using a custom mold with conically shaped needles. The mold is filled with Polyvinyl alcohol (PVA) solution and placed in a vacuum desiccator. The mold is left in the vacuum overnight until the patch is dry and ready for use. The protocol was tested with varying pressures, drying times, volume amounts, and preparation methods to determine if highquality needles can be produced. A MN is a method of DNA extraction where the patch is applied to a leaf, the needles penetrate the leaf, breaking the rigid plant cell wall to isolate the DNA. A protocol for this method of extraction was tested to ensure the patch could produce the needed yield and purity. The tests varied by the number of patches, number of applications, and plant type. To automate the MN extraction method, motion control will utilize two separate axis tables which move in the x and y directions. The y-axis table will have an end effector that fits a MN patch and will have the ability to apply the patch to the leaf sample. This end effector will also act as a lid for a downstream processing module. The other axis will position the leaf sample and processing container so that the patch can be applied accurately. The Joint Comprehensive Sequencing System (JCSS) module integrates all the components together. The output of this module feeds into the NASA Charged Information-Storage Polymer Preparation System (CHIPPS) for genomic sequencing. The extraction module operates using a series of syringes and tubing to pump the varying reagents needed for the extraction protocol. The results of the study proved that MN patches are a viable method of DNA extraction. While fabrication of high-quality needles was unsuccessful, the protocol was able to be further developed using centrifugation. The integrated design between the motion control and the JCSS enabled the potential for automation with a complete conceptual design and prototype. Future research and development for this study would include (1) further testing for fabrication (2) expanding the range of plant species compatible with the MN patch, and (3) building a working prototype for the integrated system.

Peter Ling

Functional characterization of glycosyltransferases in duckweed to enable predictive biology

Glycosyltransferases (GTs) catalyze the formation of glycosidic linkages to produce almost all complex carbohydrates. This project used a multi-disciplinary, high-throughput (HTP) biochemical and computational biology approach focused on duckweed as a model energy crop, to study carbohydrate metabolic processes. To achieve this, developed and carried out out high-throughput (HTP) functional characterization of plant glycosyltransferases (GTs) role of enzymatic microenvironments be assessed through a combined proteomic and computational biology approach, and the combined data was used to populate deep-learning frameworks to predict plant GT function. Functional validation achieved through this research is being used to assign gene function and study plant processes at the systems level to efficiently link the genome sequence with gene function. Together, the combined approaches used within this study provide a foundation for how computational prediction, in combination with high-throughput functional validation, can be used to study plant processes at the systems level and translate knowledge gained to efficiently link genome sequence with gene function in a species agnostic manner.

09 BIOMASS FUELS

Life in the Fast Lane for Protein Crystallization and X-Ray Crystallography

The common goal for structural genomic centers and consortiums is to decipher as quickly as possible the three-dimensional structures for a multitude of recombinant proteins derived from known genomic sequences. Since X-ray crystallography is the foremost method to acquire atomic resolution for macromolecules, the limiting step is obtaining protein crystals that can be useful of structure determination. High-throughput methods have been developed in recent years to clone, express, purify, crystallize and determine the three-dimensional structure of a protein gene product rapidly using automated devices, commercialized kits and consolidated protocols. However, the average number of protein structures obtained for most structural genomic groups has been very low compared to the total number of proteins purified. As more entire genomic sequences are obtained for different organisms from the three kingdoms of life, only the proteins that can be crystallized and whose structures can be obtained easily are studied. Consequently, an astonishing number of genomic proteins remain unexamined. In the era of high-throughput processes, traditional methods in molecular biology, protein chemistry and crystallization are eclipsed by automation and pipeline practices. The necessity for high rate production of protein crystals and structures has prevented the usage of more intellectual strategies and creative approaches in experimental executions. Fundamental principles and personal experiences in protein chemistry and crystallization are minimally exploited only to obtain "low-hanging fruit" protein structures. We review the practical aspects of today s high-throughput manipulations and discuss the challenges in fast pace protein crystallization and tools for crystallography. Structural genomic pipelines can be improved with information gained from low-throughput tactics that may help us reach the higher-bearing fruits. Examples of recent developments in this area are reported from the efforts of the Southeast Collaboratory for Structural Genomics (SECSG).

Pusey, Marc L.

Life in the fast lane for protein crystallization and X-ray crystallography

The common goal for structural genomic centers and consortiums is to decipher as quickly as possible the three-dimensional structures for a multitude of recombinant proteins derived from known genomic sequences. Since X-ray crystallography is the foremost method to acquire atomic resolution for macromolecules, the limiting step is obtaining protein crystals that can be useful of structure determination. High-throughput methods have been developed in recent years to clone, express, purify, crystallize and determine the three-dimensional structure of a protein gene product rapidly using automated devices, commercialized kits and consolidated protocols. However, the average number of protein structures obtained for most structural genomic groups has been very low compared to the total number of proteins purified. As more entire genomic sequences are obtained for different organisms from the three kingdoms of life, only the proteins that can be crystallized and whose structures can be obtained easily are studied. Consequently, an astonishing number of genomic proteins remain unexamined. In the era of high-throughput processes, traditional methods in molecular biology, protein chemistry and crystallization are eclipsed by automation and pipeline practices. The necessity for high-rate production of protein crystals and structures has prevented the usage of more intellectual strategies and creative approaches in experimental executions. Fundamental principles and personal experiences in protein chemistry and crystallization are minimally exploited only to obtain "low-hanging fruit" protein structures. We review the practical aspects of today's high-throughput manipulations and discuss the challenges in fast pace protein crystallization and tools for crystallography. Structural genomic pipelines can be improved with information gained from low-throughput tactics that may help us reach the higher-bearing fruits. Examples of recent developments in this area are reported from the efforts of the Southeast Collaboratory for Structural Genomics (SECSG).

Review

Revisiting synthetic lethality of Gcn5-related N-acetyltransferase (GNAT) family mutations in Haloferax volcanii

ABSTRACT Lysine acetylation is a post-translational modification that occurs in all domains of life, highlighting its evolutionary significance. Previous genome comparison identified three Gcn5-related N-acetyltransferase (GNAT) family members as lysine acetyltransferase homologs (Pat1, Pat2, and Elp3) and two deacetylase homologs (Sir2 and HdaI) in the halophilic archaeonHaloferax volcanii, withelp3andpat2proposed as a synthetic lethal gene pair. Here, we advance these findings by performing single and double mutagenesis ofelp3with thepat1andpat2lysine acetyltransferase gene homologs. Genome sequencing and PCR screens of these strains reveal successful generation of Δelp3,Δpat1Δelp3, and Δpat2Δelp3mutant strains. Although these mutant strains exhibited a reduced growth rate compared to the parent, they remained viable. Overall, this study provides genetic evidence thatelp3andpat2, while impacting cell growth, are not a synthetic lethal gene pair as previously reported. IMPORTANCE Here, we reveal by whole-genome sequencing that the GNAT family gene homologselp3andpat2can be deleted in the sameHaloferax volcaniistrain. Beyond the targeted deletions, minimal differences between the parent and Δelp3Δpat2mutant were observed, suggesting that suppressor mutations are not responsible for our ability to generate this double mutant strain. Elp3 and Pat2, thus, may not share as close a functional relationship as implied by earlier study. Our finding is significant as Elp3 is thought to function in acetylation in tRNA modification, while Pat2 likely functions in the lysine acetylation of proteins.

Microbiology

Cleanroom Microbes Survive Drying, Vacuum, and Proton Irradiation

Introduction : The goal of planetary protection at NASA is to mitigate the risk of contaminating sensitive target bodies with biological life. While many cleaning procedures have been put in place to reduce bioburden on spacecraft, microbes are experts at evolving to survive harsh conditions. Specifically, the dry, low-nutrient environment of a cleanroom (commonly used for assembly of spacecraft) can represent an environment where extremophiles can survive. Methods : Scientists at NASA MSFC wished to gather a snapshot of the microbial population within a variety of cleanrooms on site. A study was undertaken to collect air, surface, and floor samples from clean-rooms and isolate unique morphologies. From this study, 95 isolates were collected and saved in a microbial library. About 86% of these were identified at least to a genus level. Following identification, 24 microbes were selected, based on a literature review, as potential extremophiles. These were grown in liquid cultures, diluted to a set optical density, washed with water, and then applied to a sterilized Kapton coupon. Droplets were allowed to dry overnight in a biosafety cabinet. Coupons were then installed in a pelletron and pumped down to high vacuum (~1E-6 Torr). Samples were then subjected 100 keV protons at a fluence of 2x10 15 p+/cm 2 up to 4x10 15 p+/cm 2 . Following exposure, samples were returned to the microbiology lab where they were pro-cessed by submerging in water, vortexing, and then plating either droplets or spread plates. Recovery data collected was qualitative with a ranking or +, minor, or – for growth. Some selected radiotolerant strains were sequenced using the Illumina sequencing platform. The resulting genomes were annotated with the Rapid Annotations using Subsystems Technology (RAST) server and analyzed for conserved and unique stress response relevant genomic signatures to identify clues related to specific tolerances. Results and Discussion : After five rounds of proton radiation, we narrowed our isolates to five, non-spore forming bacteria that demonstrated survival: Arthrobacter koreensis, Paenarthrobacter nitroguajacolicus, Mycetocola manganoxydans , and an Erwinia sp. Furthermore, we exposed these four microbes to 254 nm wavelength light at an intensity of 80 W/m 2 at a distance of ~18 cm for 10 minutes. Only A. koreensis demonstrated survival following UV exposure. Finally, we performed whole genome sequencing on the four strains to look for genetic markers of stress resistance. When we compared the genomes of the four strains, we found that genes coding for GGDEF and EAL domains with PAS/PAC sensors were only found in A. koreensis . These domains, modulated by PAS/PAC sensors, are hypothesized to facilitate survival under drying, desiccation, and proton irradiation. Drying and Desiccation : PAS domains sense hydration changes and modulate GGDEF and EAL domain activity to adjust c-di-GMP levels, enhancing resistance to desiccation. For instance, in Pseudomonas aeruginosa , the PAS domain of RbdA modulates activity under varying hydration conditions, affecting stress responses [1]. Proton Irradiation : Proton irradiation causes oxidative stress, leading to ROS generation. PAS domains detect this stress and modulate GGDEF and EAL domains to manage oxidative stress responses. In Shewanella , EAL domain proteins modulated by PAS sensors help bacteria adapt to extreme conditions [2]. These genes upregulate other stress response genes, protecting membrane function, protein stability, DNA repair, and antioxidant defenses. The modulation of c-di-GMP by PAS domains is crucial for bacterial adaptation to stress conditions, enabling dynamic physio-logical adjustments [3]. Understanding these mechanisms provides insights into bacterial stress responses and strategies for controlling bacterial growth [4]. Conclusions : These findings indicate that clean-rooms harbor extremophile microbes that may be able to survive conditions in deep space. Furthermore, while we identified certain stress-response genes that may be at least partly responsible for the phenotypes observed in this study, there are likely unidentified genes or characteristics about A. koreensis , and other bacteria, that may allow them to survive in harsh environments. Future studies will focus on identifying these unknown genes and characteristics, further elucidating the mechanisms of extremophile survival and potentially informing the development of new biotechnologies for space exploration and other extreme environments.

Chelsi Cassilly

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat