Engineering PapersSearch

SEARCH · Engineering Papers

Results for “genome analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES

Reclassification of Botryococcus braunii chemical races into separate species based on a comparative genomics analysis

The colonial green microalga Botryococcus braunii is well known for producing liquid hydrocarbons that can be utilized as biofuel feedstocks. B. braunii is taxonomically classified as a single species made up of three chemical races, A, B, and L, that are mainly distinguished by the hydrocarbons produced. We previously reported a B race draft nuclear genome, and here we report the draft nuclear genomes for the A and L races. A comparative genomic study of the three B. braunii races and 14 other algal species within Chlorophyta revealed significant differences in the genomes of each race of B. braunii. Phylogenomically, there was a clear divergence of the three races with the A race diverging earlier than both the B and L races, and the B and L races diverging from a later common ancestor not shared by the A race. DNA repeat content analysis suggested the B race had more repeat content than the A or L races. Orthogroup analysis revealed the B. braunii races displayed more gene orthogroup diversity than three closely related Chlamydomonas species, with nearly 24-36% of all genes in each B. braunii race being specific to each race. This analysis suggests the three races are distinct species based on sufficient differences in their respective genomes. We propose reclassification of the three chemical races to the following species names: Botryococcus alkenealis (A race), Botryococcus braunii (B race), and Botryococcus lycopadienor (L race).

59 BASIC BIOLOGICAL SCIENCES

Genomic Analysis of Aspergillus Section Terrei Reveals a High Potential in Secondary Metabolite Production and Plant Biomass Degradation

Aspergillus terreus has attracted interest due to its application in industrial biotechnology, particularly for the production of itaconic acid and bioactive secondary metabolites. As related species also seem to possess a prosperous secondary metabolism, they are of high interest for genome mining and exploitation. Here, we present draft genome sequences for six species from Aspergillus section Terrei and one species from Aspergillus section Nidulantes. Whole-genome phylogeny confirmed that section Terrei is monophyletic. Genome analyses identified between 70 and 108 key secondary metabolism genes in each of the genomes of section Terrei, the highest rate found in the genus Aspergillus so far. The respective enzymes fall into 167 distinct families with most of them corresponding to potentially unique compounds or compound families. Moreover, 53% of the families were only found in a single species, which supports the suitability of species from section Terrei for further genome mining. Intriguingly, this analysis, combined with heterologous gene expression and metabolite identification, suggested that species from section Terrei use a strategy for UV protection different to other species from the genus Aspergillus. Section Terrei contains a complete plant polysaccharide degrading potential and an even higher cellulolytic potential than other Aspergilli, possibly facilitating additional applications for these species in biotechnology.

60 APPLIED LIFE SCIENCES

A genomic analysis reveals the diversity of cellulosome displaying bacteria

Introduction Several species of cellulolytic bacteria display cellulosomes, massive multi-cellulase containing complexes that degrade lignocellulosic plant biomass (LCB). A greater understanding of cellulosome structure and enzyme content could facilitate the development of new microbial-based methods to produce renewable chemicals and materials. Methods To identify novel cellulosome-displaying microbes we searched 305,693 sequenced bacterial genomes for genes encoding cellulosome proteins; dockerin-fused glycohydrolases (DocGHs) and cohesin domain containing scaffoldins. Results and discussion This analysis identified 33 bacterial species with the genomic capacity to produce cellulosomes, including 10 species not previously reported to produce these complexes, such asAcetivibrio mesophilus. Cellulosome-producing bacteria primarily originate from theAcetivibrio, Ruminococcus, Ruminiclostridium, andClostridiumgenera. A rigorous analysis of their enzyme, scaffoldin, dockerin, and cohesin content reveals phylogenetically conserved features. Based on the presence of a high number of genes encoding both scaffoldins and dockerin-fused GHs, the cellulosomes inAcetivibrioandRuminococcusbacteria possess complex architectures that are populated with a large number of distinct LCB degrading GH enzymes. Their complex cellulosomes are distinguishable by their mechanism of attachment to the cell wall, the structures of their primary scaffoldins, and by how they are transcriptionally regulated. In contrast, bacteria in theRuminiclostridiumandClostridiumgenera produce ‘simple’ cellulosomes that are constructed from only a few types of scaffoldins that based on their distinct complement of GH enzymes are predicted to exhibit high and low cellulolytic activity, respectively. Collectively, the results of this study reveal conserved and divergent architectural features in bacterial cellulosomes that could be useful in guiding ongoing efforts to harness their cellulolytic activities for bio-based chemical and materials production.

Microbiology

Unique trajectory of gene family evolution from genomic analysis of nearly all known species in an ancient yeast lineage

Gene gains and losses are a major driver of genome evolution; their precise characterization can provide insights into the origin and diversification of major lineages. Here, we examined gene family evolution of 1154 genomes from nearly all known species in the medically and technologically important yeast subphylum Saccharomycotina. We found that yeast gene family evolution differs from that of plants, animals, and filamentous ascomycetes, and is characterized by smaller overall gene numbers yet larger gene family sizes for a given gene number. Faster-evolving lineages (FELs) in yeasts experienced significantly higher rates of gene losses—commensurate with a narrowing of metabolic niche breadth—but higher speciation rates than their slower-evolving sister lineages (SELs). Gene families most often lost are those involved in mRNA splicing, carbohydrate metabolism, and cell division and are likely associated with intron loss, metabolic breadth, and non-canonical cell cycle processes. Our results highlight the significant role of gene family contractions in the evolution of yeast metabolism, genome function, and speciation, and suggest that gene family evolutionary trajectories have differed markedly across major eukaryotic lineages.

Comparative Genomics

Genomic analysis of Klebsiella aerogenes circulating in New Mexico

Klebsiella aerogenes is an opportunistic pathogen and a growing cause of healthcare-associated infections, characterized by multidrug resistance and the emergence of global high-risk clones. However, regional genomic surveillance data remain limited. Here, we sought to characterize the population structure, transmission dynamics and resistance mechanisms of clinical K. aerogenes in Albuquerque, New Mexico. We sequenced 177 clinical isolates collected between 2021 and 2023. We also developed a novel, species-specific PopPUNK database to facilitate rapid, high-resolution typing. The New Mexico K. aerogenes population was diverse but dominated by two global pandemic lineages, ST93 (47.5%) and ST4 (7.9%), which were significantly enriched for the virulence factors yersiniabactin and colibactin. Genomic evidence for recent local transmission was rare, with only four putative transmission pairs identified. The resistome was characterized by intrinsic and adaptive mutations. Nearly all isolates possessed gyrA mutations associated with decreased fluoroquinolone susceptibility. Mutations in the AmpC regulator AmpD and the outer membrane porin Omp36 were common, particularly within the dominant ST93 lineage. These mutations have been associated with increased AmpC-mediated carbapenem resistance. Our findings underscore the critical importance of genomic surveillance to monitor the transmission and evolution of adaptive resistance.

59 BASIC BIOLOGICAL SCIENCES

Genomic analysis and identification of a novel superantigen, SargEY, in Staphylococcus argenteus isolated from atopic dermatitis lesions

During surveillance of Staphylococcus aureus in lesions from patients with atopic dermatitis (AD), we isolated Staphylococcus argenteus, a species registered in 2011 as a new member of the genus Staphylococcus and previously considered a lineage of S. aureus. Genome sequence comparisons between S. argenteus isolates and representative S. aureus clinical isolates from various origins revealed that the S. argenteus genome from AD patients closely resembles that of S. aureus causing skin infections. We previously reported that 17%–22% of S. aureus isolated from skin infections produce staphylococcal enterotoxin Y (SEY), which predominantly induces T-cell proliferation via the T-cell receptor (TCR) Vα pathway. Complete genome sequencing of S. argenteus isolates revealed a gene encoding a protein similar to superantigen SEY, designated as SargEY, on its chromosome. Population structure analysis of S. argenteus revealed that these isolates are ST2250 lineage, which was the only lineage positive for the SEY-like gene among S. argenteus. Recombinant SargEY demonstrated immunological cross-reactivity with anti-SEY serum. SargEY could induce proliferation of human CD4 + and CD8 + T cells, as well as production of TNF-α and IFN-γ. SargEY showed emetic activity in a marmoset monkey model. S arg EY and SET (a phylogenetically close but uncharacterized SE) revealed their dependency on TCR Vα in inducing human T-cell proliferation. Additionally, TCR sequencing revealed other previously undescribed Vα repertoires induced by SEH. S arg EY and SEY may play roles in exacerbating the respective toxin-producing strains in AD.

59 BASIC BIOLOGICAL SCIENCES

Genomic Analysis of the Natural Variation of Fatty Acid Composition in Seed Oils of Camelina sativa

Camelina sativa is an oilseed crop that has shown strong promise as a biofuel feedstock. The profile of fatty acids greatly influences the oil quality; however, genetic mechanisms that determine the natural variation of fatty acid composition in camelina are not fully understood. A genome wide association study (GWAS) was performed to uncover genetic loci that may contribute to the contents of major fatty acids such as oleic and linolenic acids in camelina seed. Two approaches were taken to improve the GWAS efficiency. First, growing a diversity panel of 212 accessions in four locations and two nitrogen fertilization conditions revealed great variation in fatty acid contents in seeds. Second, using an improved reference genome, abundant markers, including 203,320 single nucleotide polymorphisms (SNPs) and 99,067 insertions/deletions (indels), were developed, which refined the population structure of the diversity panel. GWAS resulted in 118 genetic markers across 31 trait/treatment conditions. Closely linked markers were determined based on linkage decay and by comparing secondarily associated markers when highly associated ones were removed. Candidate genes were examined by comparing the pangenomes of 12 high-quality reference genomes. This study provides new resources to understand seed lipid metabolism and improve camelina oils through molecular breeding.

Life Sciences & Biomedicine - Other Topics

Microfabricated Genomic Analysis System

Genetic sequencing and many genetic tests and assays require electrophoretic separation of DNA. In this technique, DNA fragments are separated by size as they migrate through a sieving gel under the influence of an applied electric field. In order to conduct these analyses on-orbit, it is essential to acquire the capability to efficiently perform electrophoresis in a microgravity environment. Conventional bench top electrophoresis equipment is large and cumbersome and does not lead itself to on-orbit utilization. Much of the previous research regarding on-orbit electrophoresis involved altering conventional electrophoresis equipment for bioprocessing, purification, and/or separation technology applications. A new and more efficient approach to on-orbit electrophoresis is the use of a microfabricated electrophoresis platform. These platforms are much smaller, less expensive to produce and operate, use less power, require smaller sample sizes (nanoliters), and achieve separation in a much shorter distance (a few centimeters instead of 10 s or 100 s of centimeters.) In contrast to previous applications, this platform would be utilized as an analytical tool for life science/medical research, environmental monitoring, and medical diagnoses. Identification of infectious agents as well as radiation related damage are significant to NASA s efforts to maintain, study, and monitor crew health during and in support of near-Earth and interplanetary missions. The capability to perform genetic assays on-orbit is imperative to conduct relevant and insightful biological and medical research, as well as continuing NASA s search for life elsewhere. This technology would provide an essential analytical tool for research conducted in a microgravity environment (Shuttle, ISS, long duration/interplanetary missions.) In addition, this technology could serve as a critical and invaluable component of a biosentinel system to monitor space environment genotoxic insults to include radiation.

Gonda, Steve

Distinguishing Leptothrix and Sphaerotilus genera by an integrated genomic-phenotypic analysis supported by new Leptothrix genomes

The Sphaerotilus-Leptothrix group of bacteria includes one of the first described microorganisms, Leptothrix ochracea, an uncultured type strain, plus isolates of Leptothrix and Sphaerotilus. This group is unified by the ability to form sheaths and oxidize metals, although L. ochracea exhibits obvious ecological, morphological, and functional differences from the rest of Sphaerotilus-Leptothrix. Recently, there have been calls to combine the group into one genus, Sphaerotilus; however, these studies lacked adequate genomic representation of L. ochracea. Here, we present a comprehensive comparative genomic analysis of the Sphaerotilus-Leptothrix group, including expanded representation of L. ochracea, a closely related novel species, Leptothrix toolikensis, and two new isolates (Leptothrix mechoopdaensis). Analysis of 38 genomes resolves three phylogenetic and functional groups: the ochracea-type Leptothrix (Group 1), the mobilis-type Leptothrix (Group 2), and Sphaerotilus (Group 3). Group 1 genomes form a separate genus based on average nucleotide identity and alignment fraction. The genomes clearly diverge from the rest of Sphaerotilus-Leptothrix in phylogeny, size, and metabolic potential. Group 1 genomes are much smaller (2.59–3.04 Mb) than those of Groups 2 (4.55–6.06 Mb) and 3 (3.94–5.07 Mb), while encoding more metal oxidases and fewer carbohydrate-active enzymes. Group 2 clusters with Group 3 phylogenetically and is similar in organic carbon metabolisms but maintains more metal oxidation genes. Group 2 members lack homogeneity in phenotype and genotype, suggesting that additional isolates and genomes are needed for confident classification. However, Group 1 genomes (L. ochracea and L. toolikensis) show clear divergence, precluding their inclusion in Sphaerotilus and supporting the retention of the genus Leptothrix.

Leptothrix

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology

Analysis of genomic signatures associated with Variovorax endosphere colonization

This repository contains the analysis code and supporting datasets associated with the study “Genomic signatures in Variovorax enabling colonization of the Populus endosphere.” Beals DG, Carper DL, Hochanadel LH, Jawdy SS, Klingeman DM, Piatkowski BT, Weston DJ, Doktycz MJ, Pelletier DA. 2026. Genomic signatures in Variovorax enabling colonization of the Populus endosphere. mSystems 11:e01605-25. https://doi.org/10.1128/msystems.01605-25 The scripts are organized sequentially (01–07) and document the workflows used for: Sequence-read alignment and feature counting Orthogroup and KEGG Ortholog annotation Count normalization Statistical analysis and aggregation Generation of manuscript figures and tables Repository contents The uncompressed files are the finalized, formatted datasets used to generate the figures and tables reported in the study, including the supplemental CSV files referenced in the manuscript. The accompanying ZIP archive contains the complete codebase and example data_input/ and data_output/ directories illustrating the organization and execution of the analytical workflow. Individual scripts identify the corresponding manuscript analyses and figure panels. Raw sequencing data Raw sequencing reads are available through the NCBI Sequence Read Archive under BioProject accession PRJNA1322484.

Beals, Delaney [ORNL] (ORCID:0000000306274574)

Methanococcus jannaschii genome: revisited

Analysis of genomic sequences is necessarily an ongoing process. Initial gene assignments tend (wisely) to be on the conservative side (Venter, 1996). The analysis of the genome then grows in an iterative fashion as additional data and more sophisticated algorithms are brought to bear on the data. The present report is an emendation of the original gene list of Methanococcus jannaschii (Bult et al., 1996). By using a somewhat more updated database and more relaxed (and operator-intensive) pattern matching methods, we were able to add significantly to, and in a few cases amend, the gene identification table originally published by Bult et al. (1996).

Non-NASA Center

Genome-resolved analysis of Serratia marcescens strain SMTT infers niche specialization as a hydrocarbon-degrader

Abstract Bacteria that are chronically exposed to high levels of pollutants demonstrate genomic and corresponding metabolic diversity that complement their strategies for adaptation to hydrocarbon-rich environments. Whole genome sequencing was carried out to infer functional traits of Serratia marcescens strain SMTT recovered from soil contaminated with crude oil. The genome size (Mb) was 5,013,981 with a total gene count of 4,842. Comparative analyses with carefully selected S. marcescens strains, 2 of which are associated with contaminated soil, show conservation of central metabolic pathways in addition to intra-specific genetic diversity and metabolic flexibility. Genome comparisons also indicated an enrichment of genes associated with multidrug resistance and efflux pumps for SMTT. The SMTT genome contained genes that enable the catabolism of aromatic compounds via the protocatechuate para-degradation pathway, in addition to meta-cleavage of catechol (meta-cleavage pathway II); gene enrichment for aromatic compound degradation was markedly higher for SMTT compared to the other S. marcescens strains analysed. Our data presents a valuable genetic inventory for future studies on strains of S. marcescens and provides insights into those genomic features of SMTT with industrial potential.

Genetics & Heredity

Genome-scale analysis of interactions between genetic perturbations and natural variation

Interactions between genetic perturbations and segregating loci can cause perturbations to show different phenotypic effects across genetically distinct individuals. To study these interactions on a genome scale in many individuals, we used combinatorial DNA barcode sequencing to measure the fitness effects of 8046 CRISPRi perturbations targeting 1721 distinct genes in 169 yeast cross progeny (or segregants). We identified 460 genes whose perturbation has different effects across segregants. Several factors caused perturbations to show variable effects, including baseline segregant fitness, the mean effect of a perturbation across segregants, and interacting loci. We mapped 234 interacting loci and found four hub loci that interact with many different perturbations. Perturbations that interact with a given hub exhibit similar epistatic relationships with the hub and show enrichment for cellular processes that may mediate these interactions. These results suggest that an individual’s response to perturbations is shaped by a network of perturbation-locus interactions that cannot be measured by approaches that examine perturbations or natural variation alone.

59 BASIC BIOLOGICAL SCIENCES

Assessing the Application of a Genomic Network Analysis in Population Ecology: Inferring Patterns of Dispersal and Geographic Structure in the Emerging Pathogen, Coccidioides

A challenge in population ecology studies is identifying how to best group individuals into populations, especially when individual origin is unknown. Machine learning has improved upon traditional methods of identifying population structure and is more efficient at handling large, complex datasets. We demonstrate the applicability of a machine learning method to identify hierarchical population structure in an emerging pathogen, Coccidioides spp., the causative agent of Valley fever. We compared the network clusters to structure identified by traditional tools as a validation of the network performance. We used publicly available whole-genome data for 48 C. immitis and 102 C. posadasii, resulting in 168,211 genome-wide SNPs among the two species. The network analysis grouped samples into populations comparable to the literature for these species but also identified fine-scale geographic structure and travel-associated cases not reported thus far. Exploring different resolutions in the network made it easy to identify unique genotypes specific to California and possibly Nevada, as well as Phoenix- and Tucson-acquired infections in non-endemic areas, regardless of reported travel history. The present study provides a promising example of how a ML-based network analysis can improve our ability to understand pathogen ecology, group cases into populations and infer travel-associated infections.

59 BASIC BIOLOGICAL SCIENCES

Aeromonas in South Asia: genomic insights into an environmental pathogen and reservoir of antimicrobial resistance

Aeromonads are an ecologically versatile group of bacteria that cause infections in aquatic animals and are recognised as emerging human pathogens. Despite this, our understanding of Aeromonas diversity, especially the relationship between clinical and environmental strains, remains limited. Here, we present a genomic analysis of the Aeromonas genus, comprising 1853 genomes, and a detailed comparison of clinical and environmental strains from South Asia, including 996 newly sequenced genomes from Bangladesh and India. Phylogenetic analyses revealed that Aeromonas is a highly diverse genus, with no distinct clade separating clinical and environmental isolates. We identified 28 Aeromonas species and 905 novel sequence types, comprising 72.5% of the genomes. Notably, we show a high incidence of antimicrobial resistance (AMR) genes across all isolates, including against front and last-line antibiotics. Finally, we highlight frequent misidentification of Aeromonas as Vibrio cholerae, which is relevant to cholera-endemic regions where both genera co-exist and are associated with diarrhoeal disease. Our study underscores Aeromonas as an important environmental AMR reservoir and emerging multi-species pathogen capable of spilling over into human populations.

59 BASIC BIOLOGICAL SCIENCES