Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Ecological Trait-Based Digital Categorization of Microbial Genomes for Denitrification Potential

Microorganisms encode proteins that function in the transformations of useful and harmful nitrogenous compounds in the global nitrogen cycle. The major transformations in the nitrogen cycle are nitrogen fixation, nitrification, denitrification, anaerobic ammonium oxidation, and ammonification. The focus of this report is the complex biogeochemical process of denitrification, which, in the complete form, consists of a series of four enzyme-catalyzed reduction reactions that transforms nitrate to nitrogen gas. Denitrification is a microbial strain-level ecological trait (characteristic), and denitrification potential (functional performance) can be inferred from trait rules that rely on the presence or absence of genes for denitrifying enzymes in microbial genomes. Despite the global significance of denitrification and associated large-scale genomic and scholarly data sources, there is lack of datasets and interactive computational tools for investigating microbial genomes according to denitrification trait rules. Therefore, our goal is to categorize archaeal and bacterial genomes by denitrification potential based on denitrification traits defined by rules of enzyme involvement in the denitrification reduction steps. We report the integration of datasets on genome, taxonomic lineage, ecosystem, and denitrifying enzymes to provide data investigations context for the denitrification potential of microbial strains. We constructed an ecosystem and taxonomic annotated denitrification potential dataset of 62,624 microbial genomes (866 archaea and 61,758 bacteria) that encode at least one of the twelve denitrifying enzymes in the four-step canonical denitrification pathway. Our four-digit binary-coding scheme categorized the microbial genomes to one of sixteen denitrification traits including complete denitrification traits assigned to 3280 genomes from 260 bacteria genera. The bacterial strains with complete denitrification potential pattern included Arcobacteraceae strains isolated or detected in diverse ecosystems including aquatic, human, plant, and Mollusca (shellfish). The dataset on microbial denitrification potential and associated interactive data investigations tools can serve as research resources for understanding the biochemical, molecular, and physiological aspects of microbial denitrification, among others. The microbial denitrification data resources produced in our research can also be useful for identifying microbial strains for synthetic denitrifying communities.

59 BASIC BIOLOGICAL SCIENCES↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

A fast comparative genome browser for diverse bacteria and archaea

Genome sequencing has revealed an incredible diversity of bacteria and archaea, but there are no fast and convenient tools for browsing across these genomes. It is cumbersome to view the prevalence of homologs for a protein of interest, or the gene neighborhoods of those homologs, across the diversity of the prokaryotes. We developed a web-based tool, fast . genomics , that uses two strategies to support fast browsing across the diversity of prokaryotes. First, the database of genomes is split up. The main database contains one representative from each of the 6,377 genera that have a high-quality genome, and additional databases for each taxonomic order contain up to 10 representatives of each species. Second, homologs of proteins of interest are identified quickly by using accelerated searches, usually in a few seconds. Once homologs are identified, fast . genomics can quickly show their prevalence across taxa, view their neighboring genes, or compare the prevalence of two different proteins. Fast . genomics is available at https://fast.genomics.lbl.gov .

59 BASIC BIOLOGICAL SCIENCES↗

Make the Most Out of Genome Announcements with KBase

Genomics is a dynamic field driven by new technology, massive growth in data, and changes in how scientists share knowledge. This brings new challenges for researchers looking to maximize the impact of their science and ensure the longevity of their data. These changes are especially relevant for the practice of publishing genome announcements. Genome announcements are meant to share and promote newly sequenced organisms with the larger scientific community. A few decades ago, genome announcements were highly anticipated hallmarks of novel science. But, what was once a publishing coup is now an every day event. It is harder and harder to keep abreast of all the genomes that have been sequenced and promote your own new genomes to get the attention and recognition of your research community. Fortunately, KBase can help construct genomes from your sequencing data, generate reports for your announcement, and promote your data to your research community.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic genome evolution in a model fern

The large size and complexity of most fern genomes have hampered efforts to elucidate fundamental aspects of fern biology and land plant evolution through genome-enabled research. Here we present a chromosomal genome assembly and associated methylome, transcriptome and metabolome analyses for the model fern species Ceratopteris richardii. The assembly reveals a history of remarkably dynamic genome evolution including rapid changes in genome content and structure following the most recent whole-genome duplication approximately 60 million years ago. These changes include massive gene loss, rampant tandem duplications and multiple horizontal gene transfers from bacteria, contributing to the diversification of defence-related gene families. The insertion of transposable elements into introns has led to the large size of the Ceratopteris genome and to exceptionally long genes relative to other plants. Gene family analyses indicate that genes directing seed development were co-opted from those controlling the development of fern sporangia, providing insights into seed plant evolution. Our findings and annotated genome assembly extend the utility of Ceratopteris as a model for investigating and teaching plant biology.

59 BASIC BIOLOGICAL SCIENCES↗

Population structure limits the use of genomic data for predicting phenotypes and managing genetic resources in forest trees

There is overwhelming evidence that forest trees are locally adapted to climate. Thus, genecological models based on population phenotypes have been used to measure local adaptation, infer genetic maladaptation to climate, and guide assisted migration. However, instead of phenotypes, there is increasing interest in using genomic data for gene resource management. We used whole-genome resequencing and common-garden experiments to understand the genetic architecture of adaptive traits in black cottonwood. We studied the potential of using genome-wide association studies (GWAS) and genomic prediction to detect causal loci, identify climate-adapted phenotypes, and inform gene resource management. We analyzed population structure by partitioning phenotypic and genomic (single-nucleotide polymorphism) variation among 840 genotypes collected from 91 stands along 16 rivers. Most phenotypic variation (60 to 81%) occurred among populations and was strongly associated with climate. Population phenotypes were predicted well using genomic data (e.g., predictive abilityr> 0.9) but almost as well using climate or geography (r> 0.8). In contrast, genomic prediction within populations was poor (r< 0.2). We identified many GWAS associations among populations, but most appeared to be spurious based on pooled within-population analyses. Hierarchical partitioning of linkage disequilibrium and haplotype sharing suggested that within-population genomic prediction and GWAS were poor because allele frequencies of causal loci and linked markers differed among populations. Given the urgent need to conserve natural populations and ecosystems, our results suggest that climate variables alone can be used to predict population phenotypes, delineate seed zones and deployment zones, and guide assisted migration.

Science & Technology - Other Topics↗

Improved genome editing by an engineered CRISPR-Cas12a

CRISPR-Cas12a is an RNA-guided, programmable genome editing enzyme found within bacterial adaptive immune pathways. Unlike CRISPR-Cas9, Cas12a uses only a single catalytic site to both cleave target double-stranded DNA (dsDNA) (cis-activity) and indiscriminately degrade single-stranded DNA (ssDNA) (trans-activity). To investigate how the relative potency of cis- versus trans-DNase activity affects Cas12a-mediated genome editing, we first used structure-guided engineering to generate variants of Lachnospiraceae bacterium Cas12a that selectively disrupt trans-activity. The resulting engineered mutant with the biggest differential between cis - and trans -DNase activity in vitro showed minimal genome editing activity in human cells, motivating a second set of experiments using directed evolution to generate additional mutants with robust genome editing activity. Notably, these engineered and evolved mutants had enhanced ability to induce homology-directed repair (HDR) editing by 2–18-fold compared to wild-type Cas12a when using HDR donors containing mismatches with crRNA at the PAM-distal region. Finally, a site-specific reversion mutation produced improved Cas12a (iCas12a) variants with superior genome editing efficiency at genomic sites that are difficult to edit using wild-type Cas12a. This strategy establishes a pipeline for creating improved genome editing tools by combining structural insights with randomization and selection. The available structures of other CRISPR-Cas enzymes will enable this strategy to be applied to improve the efficacy of other genome-editing proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Parallel String Graph Construction and Transitive Reduction for De Novo Genome Assembly

One of the most computationally intensive tasks in computational biology is de novo genome assembly, the decoding of the sequence of an unknown genome from redundant and erroneous short sequences. A common assembly paradigm identifies overlapping sequences, simplifies their layout, and creates consensus. Despite many algorithms developed in the literature, the efficient assembly of large genomes is still an open problem. In this work, we introduce new distributed-memory parallel algorithms for overlap detection and layout simplification steps of de novo genome assembly, and implement them in the diBELLA 2D pipeline. Our distributed memory algorithms for both overlap detection and layout simplification are based on linear-algebra operations over semirings using 2D distributed sparse matrices. Our layout step consists of performing a transitive reduction from the overlap graph to a string graph. We provide a detailed communication analysis of the main stages of our new algorithms. diBELLA 2D achieves near linear scaling with over 80% parallel efficiency for the human genome, reducing the runtime for overlap detection by 1.2-1.3× for the human genome and 1.5-1.9× for C.elegans compared to the state-of-the-art. Our transitive reduction algorithm outperforms an existing distributed-memory implementation by 10.5-13.3× for the human genome and 18-29× for the C. elegans. Our work paves the way for efficient de novo assembly of large genomes using long reads in distributed memory.

59 BASIC BIOLOGICAL SCIENCES↗

CYPminer: an automated cytochrome P450 identification, classification, and data analysis tool for genome data sets across kingdoms

Background: Cytochrome P450 monooxygenases (termed CYPs or P450s) are hemoproteins ubiquitously found across all kingdoms, playing a central role in intracellular metabolism, especially in metabolism of drugs and xenobiotics. The explosive growth of genome sequencing brings a new set of challenges and issues for researchers, such as a systematic investigation of CYPs across all kingdoms in terms of identification, classification, and pan-CYPome analyses. Such investigation requires an automated tool that can handle an enormous amount of sequencing data in a timely manner. Results: CYPminer was developed in the Python language to facilitate rapid, comprehensive analysis of CYPs from genomes of all kingdoms. CYPminer consists of two procedures i) to generate the Genome-CYP Matrix (GCM) that lists all occurrences of CYPs across the genomes, and ii) to perform analyses and visualization of the GCM, including pan-CYPomes (pan- and core-CYPome), CYP co-occurrence networks, CYP clouds, and genome clustering data. The performance of CYPminer was evaluated with three datasets from fungal and bacterial genome sequences. Conclusions: CYPminer completes CYP analyses for large-scale genomes from all kingdoms, which allows systematic genome annotation and comparative insights for CYPs. CYPminer also can be extended and adapted easily for broader usage.

59 BASIC BIOLOGICAL SCIENCES↗

Importance of appropriate genome information for the design of mating type primers in black and yellow morel populations

Abstract Morels are highly prized edible fungi where sexual reproduction is essential for fruiting-body production. As a result, a comprehensive understanding of their sexual reproduction is of great interest. Central to this is the identification of the reproductive strategies used by morels. Sexual reproduction in fungi is controlled by mating-type ( MAT ) genes and morels are thought to be mainly heterothallic with two idiomorphs, MAT1-1 and MAT1-2. Genomic sequencing of black (Elata clade) and yellow (Esculenta clade) morel species has led to the development of PCR primers designed to amplify genes from the two idiomorphs for rapid genotyping of isolates from these two clades. To evaluate the design and theoretical performance of these primers we performed a thorough bioinformatic investigation, including the detection of the MAT region in publicly available Morchella genomes and in-silico PCR analyses. All examined genomes, including those used for primer design, appeared to be heterothallic. This indicates an inherent fault in the original primer design which utilized a single Morchella genome, as the use of two genomes with complementary mating types would be required to design accurate primers for both idiomorphs. Furthermore, potential off-targets were identified for some of the previously published primer sets, but verification was challenging due to lack of adequate genomic information and detailed methodologies for primer design. Examinations of the black morel specific primer pairs (MAT11L/R and MAT22L/R) indicated the MAT22 primers would correctly target and amplify the MAT1-2 idiomorph, but the MAT11 primers appear to be capable of amplifying incorrect off-targets within the genome. The yellow morel primer pairs (EMAT1-1 L/R and EMAT1-2 L/R) appear to have reporting errors, as the published primer sequences are dissimilar with reported amplicon sequences and the EMAT1-2 primers appear to amplify the RNA polymerase II subunit ( RPB2 ) gene. The lack of the reference genome used in primer design and descriptive methodology made it challenging to fully assess the apparent issues with the primers for this clade. In conclusion, additional work is still required for the generation of reliable primers to investigate mating types in morels and to assess their performance on different clades and across multiple geographical regions.

59 BASIC BIOLOGICAL SCIENCES↗

IMA genome-F18

Sequencing fungal genomes has now become very common and the list of genomes in this manuscript reflects this. Particularly relevant is that the first announcement is a re-identification of Penicillium genomes available on NCBI. The fact that more than 100 of these genomes have been deposited without the correct species names speak volumes to the fact that we must continue training fungal taxonomists and the importance of the International Mycological Association (after which this journal is named). When we started the genome series in 2013, one of the essential aspects was the need to have a phylogenetic tree as part of the manuscript. This came about as the result of a discussion with colleagues in NCBI who were trying to deal with the very many incorrectly identified bacterial genomes (at the time) which had been submitted to NCBI. We are now in the same position with fungal genomes. Sequencing a fungal genome is all too easy but providing a correct species name and ensuring that the fungus has in fact been correctly identified seems to be more difficult. We know that there are thousands of fungi which have not yet been described. The availability of sequence data has made identification of fungi easier but also serves to highlight the need to have a fungal taxonomist in the project to make sure that mistakes are not made.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid identification of enteric bacteria from whole genome sequences using average nucleotide identity metrics

Identification of enteric bacteria species by whole genome sequence (WGS) analysis requires a rapid and an easily standardized approach. We leveraged the principles of average nucleotide identity using MUMmer (ANIm) software, which calculates the percent bases aligned between two bacterial genomes and their corresponding ANI values, to set threshold values for determining species consistent with the conventional identification methods of known species. The performance of species identification was evaluated using two datasets: the Reference Genome Dataset v2 (RGDv2), consisting of 43 enteric genome assemblies representing 32 species, and the Test Genome Dataset (TGDv1), comprising 454 genome assemblies which is designed to represent all species needed to query for identification, as well as rare and closely related species. The RGDv2 contains six Campylobacter spp., three Escherichia/Shigella spp., one Grimontia hollisae, six Listeria spp., one Photobacterium damselae, two Salmonella spp., and thirteen Vibrio spp., while the TGDv1 contains 454 enteric bacterial genomes representing 42 different species. The analysis showed that, when a standard minimum of 70% genome bases alignment existed, the ANI threshold values determined for these species were ≥95 for Escherichia/Shigella and Vibrio species, ≥93% for Salmonella species, and ≥92% for Campylobacter and Listeria species. Using these metrics, the RGDv2 accurately classified all validation strains in TGDv1 at the species level, which is consistent with the classification based on previous gold standard methods.

59 BASIC BIOLOGICAL SCIENCES↗

A haplotype-resolved, chromosome-scale genome assembly for the southern live oak, Quercus virginiana

Hybridization is a major force driving diversification, migration, and adaptation in Quercus species. While population genetics and phylogenetics have traditionally been used for studying these processes, advances in sequencing technology now enable us to incorporate comparative and pan-genomic approaches as well. Here, we present a highly contiguous, chromosome-scale and haplotype-resolved genome assembly for the southern live oak, Quercus virginiana, the first reference genome for section Virentes, as part of the American Campus Tree Genomes program. Originating from a clone of Auburn University's historic “Toomer's Oak,” this assembly contributes to the pool of genomic resources for investigating recombination, haplotype variation, and structural genomic changes influencing hybridization potential in this clade and across Quercus. It also provides insights into the architecture of the putative centromeric regions within the genus. Alongside other oak references, the Q. virginiana genome will support research into the evolution and adaptation of the Quercus genus.

Quercus virginiana↗

The genetics of aerotolerant growth in an alphaproteobacterium with a naturally reduced genome

Reduced genome bacteria are genetically simplified systems that facilitate biological study and industrial use. The free-living alphaproteobacterium Zymomonas mobilis has a naturally reduced genome containing fewer than 2,000 protein-coding genes. Despite its small genome, Z. mobilis thrives in diverse conditions including the presence or absence of atmospheric oxygen. However, insufficient characterization of essential and conditionally essential genes has limited broader adoption of Z. mobilis as a model alphaproteobacterium. Here, we use genome-scale CRISPRi-seq (clustered regularly interspaced short palindromic repeats interference sequencing) to systematically identify and characterize Z. mobilis genes that are conditionally essential for aerotolerant or anaerobic growth or are generally essential across both conditions. Comparative genomics revealed that the essentiality of most “generally essential” genes was shared between Z. mobilis and other Alphaproteobacteria, validating Z. mobilis as a reduced genome model. Among conditionally essential genes, we found that the DNA repair gene, recJ, was critical only for aerobic growth but reduced the mutation rate under both conditions. Further, we show that genes encoding the F1FO ATP synthase and Rhodobacter nitrogen fixation (Rnf) respiratory complex are required for the anaerobic growth of Z. mobilis. Combining CRISPRi partial knockdowns with metabolomics and membrane potential measurements, we determined that the ATP synthase generates membrane potential that is consumed by Rnf to power downstream processes. Rnf knockdown strains accumulated isoprenoid biosynthesis intermediates, suggesting a key role for Rnf in powering essential biosynthetic reactions. Our work establishes Z. mobilis as a streamlined model for alphaproteobacterial genetics, has broad implications in bacterial energy coupling, and informs Z. mobilis genome manipulation for optimized production of valuable isoprenoid-based bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomics reveals the high diversity and adaptation strategies of Polaromonas from polar environments

Abstract Background Bacteria from the genus Polaromonas are dominant phylotypes found in a variety of low-temperature environments in polar regions. The diversity and biogeographic distribution of Polaromonas have been largely expanded on the basis of 16 S rRNA gene amplicon sequencing. However, the evolution and cold adaptation mechanisms of Polaromonas from polar regions are poorly understood at the genomic level. Results A total of 202 genomes of the genus Polaromonas were analyzed, and 121 different species were delineated on the basis of average nucleotide identity (ANI) and phylogenomic placements. Remarkably, 8 genomes recovered from polar environments clustered into a separate clade (‘polar group’ hereafter). The genome size, coding density and coding sequences (CDSs) of the polar group were significantly different from those of other nonpolar Polaromonas . Furthermore, the enrichment of genes involved in carbohydrate and peptide metabolism was evident in the polar group. In addition, genes encoding proteins related to betaine synthesis and transport were increased in the genomes from the polar group. Phylogenomic analysis revealed that two different evolutionary scenarios may explain the adaptation of Polaromonas to cold environments in polar regions. Conclusions The global distribution of the genus Polaromonas highlights its strong adaptability in both polar and nonpolar environments. Species delineation significantly expands our understanding of the diversity of the Polaromonas genus on a global scale. In this study, a polar-specific clade was found, which may represent a specific ecotype well adapted to polar environments. Collectively, genomic insight into the metabolic diversity, evolution and adaptation of the genus Polaromonas at the genome level provides a genetic basis for understanding the potential response mechanisms of Polaromonas to global warming in polar regions.

54 ENVIRONMENTAL SCIENCES↗

The architecture of resilience: a genome assembly of Myrothamnus flabellifolia sheds light on desiccation tolerance and sex determination

Myrothamnus flabellifolia is a dioecious resurrection plant endemic to southern Africa that has become an important model for understanding desiccation tolerance. Despite its ecological and medicinal significance, genomic and transcriptomic resources for the species are limited. We generated a chromosome-level, haplotype-resolved reference genome assembly and annotation for M. flabellifolia and conducted transcriptomic profiling across a natural dehydration–rehydration time course in the field. Genome architecture and sex determination were characterized, and co-expression network and cis-regulatory element (CRE) enrichment analyses were used to investigate dynamic responses to desiccation. The 1.28-Gb genome exhibits unusually consistent chromatin architecture with unique chromosome organization across highly divergent haplotypes. We identified an XY sexual system with a small sex-determining region on Chromosome 8. Transcriptomic responses varied with dehydration severity, pointing to early suppression of growth, progressive activation of protective mechanisms, and subsequent return to homeostasis upon rehydration. Late embryogenesis abundant and early light-induced protein transcripts were dynamically regulated and showed enrichment of abscisic acid and stress-responsive CREs pointing toward conserved responses. Together, this study provides foundational resources for understanding the genomic architecture and reproductive biology of M. flabellifolia and offers new insights into the mechanisms of desiccation tolerance.

chromosome structure↗

Edaphic controls on genome size and GC content of bacteria in soil microbial communities

Nutrient limitation has been shown to reduce bacterial genome size and influence nucleotide composition; however, much of this work has been conducted in marine systems and the factors which shape soil bacterial genomic traits remain largely unknown. Here, for this work, we determined average genome size, GC content, codon usage, and amino acid content from 398 soil metagenomes across a broad geographic range and used machine-learning to determine the environmental parameters that most strongly explain the distribution of these traits. We found that genomic trait averages were most related to pH, which we suggest is primarily due to the correlation of pH with several environmental parameters, particularly soil carbon content. Low pH soils had higher carbon to nitrogen ratios (C:N) and tended to have communities with lower GC content and larger genomes, potentially a response to increased physiological stress and a requirement for metabolic diversity. Conversely, communities in high pH and low soil C:N had smaller genomes and higher GC content—indicating potential resource driven selection against AT base pairs, which have a higher C:N than GC base pairs. Similarly, we found that nutrient conservation also applied to amino acid stoichiometry, where bacteria in soils with low C:N ratios tended to code for amino acids with lower C:N. Together, these relationships point towards fundamental mechanisms that underpin genome size, and nucleotide and amino acid selection in soil bacteria.

54 ENVIRONMENTAL SCIENCES↗