Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Creation of an Acyltransferase Toolbox for Plant Biomass Engineering (Final Report)

The major goal of this project was to expand our understanding of acyl‐CoA ligases and BAHD acyltransferases and their utility in plant engineering. We combined bioinformatic analysis of genes and transcripts with functional fingerprinting of synthesized genes produced by JGI. Best candidates from this experimental pipeline were transferred into bioenergy plants to study their effects on lignin composition. We found combinations of ligase and transferase genes encoding enzymes with interesting catalytic specificities. Our work demonstrated the feasibility of use of acyl-CoA ligases and BAHD acyltransferases to alter the composition of plant cell walls without deleterious effects on the modified plant.

59 BASIC BIOLOGICAL SCIENCES↗

Case Study: Can you find Delftia?

Delftia is a genus of bacteria with a bunch of cool features! The best studied species, Delftia acidovorans can produce gold nanoparticles from gold ions in solution. This bacteria has been found living in biofolms with Cupriavidis metallidurans on gold nuggets. It is also found in soil, in sinks and in rhizospheres of different plants where it promotes their growth. Delftia acidovorans forms gold nuggets by producing a short nonribosomal peptide called delftibactin. The 16 genes responsible for the production of delftibactin are called the del cluster (delA-delP). These genes were originally discovered in Delftia acidovorans SPH-1, but our research shows that the del cluster appears to be present across the genus. We'll start with a number of raw reads, clean them up and do a little taxonomy to figure out what we might be able to assemble. Next, we'll assemble the reads using 4 different methods and pick out the best assembly. Then we'll take those reads and sort them by genome and assess the quality of those genomes. We'll select the high quality genomes to annotate and insert into phylogenetic trees to find relatives.

59 BASIC BIOLOGICAL SCIENCES↗

METABOLIC: high-throughput profiling of microbial genomes for functional traits, metabolism, biogeochemistry, and community-scale functional networks

Background Advances in microbiome science are being driven in large part due to our ability to study and infer microbial ecology from genomes reconstructed from mixed microbial communities using metagenomics and single-cell genomics. Such omics-based techniques allow us to read genomic blueprints of microorganisms, decipher their functional capacities and activities, and reconstruct their roles in biogeochemical processes. Currently available tools for analyses of genomic data can annotate and depict metabolic functions to some extent; however, no standardized approaches are currently available for the comprehensive characterization of metabolic predictions, metabolite exchanges, microbial interactions, and microbial contributions to biogeochemical cycling. Results We present METABOLIC (METabolic And BiogeOchemistry anaLyses In miCrobes), a scalable software to advance microbial ecology and biogeochemistry studies using genomes at the resolution of individual organisms and/or microbial communities. The genome-scale workflow includes annotation of microbial genomes, motif validation of biochemically validated conserved protein residues, metabolic pathway analyses, and calculation of contributions to individual biogeochemical transformations and cycles. The community-scale workflow supplements genome-scale analyses with determination of genome abundance in the microbiome, potential microbial metabolic handoffs and metabolite exchange, reconstruction of functional networks, and determination of microbial contributions to biogeochemical cycles. METABOLIC can take input genomes from isolates, metagenome-assembled genomes, or single-cell genomes. Results are presented in the form of tables for metabolism and a variety of visualizations including biogeochemical cycling potential, representation of sequential metabolic transformations, community-scale microbial functional networks using a newly defined metric “MW-score” (metabolic weight score), and metabolic Sankey diagrams. METABOLIC takes ~ 3 h with 40 CPU threads to process ~ 100 genomes and corresponding metagenomic reads within which the most compute-demanding part of hmmsearch takes ~ 45 min, while it takes ~ 5 h to complete hmmsearch for ~ 3600 genomes. Tests of accuracy, robustness, and consistency suggest METABOLIC provides better performance compared to other software and online servers. To highlight the utility and versatility of METABOLIC, we demonstrate its capabilities on diverse metagenomic datasets from the marine subsurface, terrestrial subsurface, meadow soil, deep sea, freshwater lakes, wastewater, and the human gut. Conclusion METABOLIC enables the consistent and reproducible study of microbial community ecology and biogeochemistry using a foundation of genome-informed microbial metabolism, and will advance the integration of uncultivated organisms into metabolic and biogeochemical models. METABOLIC is written in Perl and R and is freely available under GPLv3 at https://github.com/AnantharamanLab/METABOLIC.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-Scale Transcription-Translation Mapping Reveals Features of Zymomonas mobilis Transcription Units and Promoters

Efforts to rationally engineer synthetic pathways in Zymomonas mobilis are impeded by a lack of knowledge and tools for predictable and quantitative programming of gene regulation at the transcriptional, posttranscriptional, and posttranslational levels. With the detailed functional characterization of the Z. mobilis genome presented in this work, we provide crucial knowledge for the development of synthetic genetic parts tailored to Z. mobilis . This information is vital as researchers continue to develop Z. mobilis for synthetic biology applications. Our methods and statistical analyses also provide ways to rapidly advance the understanding of poorly characterized bacteria via empirical data that enable the experimental validation of sequence-based prediction for genome characterization and annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Draft genome of multiple resistance donor plant Sinapis alba: An insight into SSRs, annotations and phylogenetics

Sinapis alba is a wild member of the Brassicaceae family reported to possess genetic resistance against major biotic and abiotic stresses of oilseed brassicas. However, the resistance nature of S. alba was not exploited generously due to the unavailability of usable genome sequences in public databases. Therefore, the present study was conducted to assemble the first draft genome from raw whole genome shotgun sequences with annotation and develop simple sequence repeat markers for molecular genetics and marker-assisted breeding. Results The raw genome sequences had 96x coverage on the Illumina platform with 170 Gbp data. The developed assembly by SOAPdenovo2 has ~459 Mbp genome size covered in 403,423 contigs with an average size of 1138.04 bp. The assembly was BLASTX with Arabidopsis thaliana which showed 32.9% positive hits between both plants. The top hit species distribution analysis showed the highest similarity with A. thaliana. A total of 809,597 GO level annotations were recorded after BLASTX results, and 34,012 sequences were annotated with different enzyme codes grouped under seven classes. The gene prediction tool AUGUSTUS identified 113,107 probable genes with an average size of 684 bp. The biochemical pathway annotation assigned 16,119 potential genes to 152 KEGG maps and 1751 enzyme codes. The development of potential SSRs from the de-novo assembly yielded 70731 unique primer pairs. Out of 159 randomly selected SSR markers for validation, 149 successfully amplified in S. alba. However, 10 SSR markers did not amplify during the validation experiment. Conclusion The annotated genome assembly with a large number of SSRs was developed in the present study. To the best of our knowledge, this is the first report of S. alba genome assembly development, annotation, and SSRs mining to date. The data presented here will be a very important resource for future crop improvement programs, especially for resistant breeding.

59 BASIC BIOLOGICAL SCIENCES↗

A multi-omic characterization of the physiological responses to salt stress in Scenedesmus obliquus UTEX393

Scenedesmus obliquus UTEX393 is a promising microalgal candidate for sustainable biomanufacturing but its limited halotolerance hinders large-scale cultivation in saline environments. To investigate the molecular basis of salt stress responses, we conducted a comprehensive multi-omic analysis integrating genomics, transcriptomics, proteomics, lipidomics, metabolomics, and DNA affinity purification sequencing (DAP-seq). An improved nuclear genome assembly and annotation yielded 19,017 gene models and a 97% BUSCO completeness score, enabling construction of a genome-scale metabolic model. Comparing 15 ppt salinity stress to 5 ppt control, growth and productivity were significantly reduced, accompanied by widespread transcriptomic and proteomic changes. Transcriptomic analysis revealed downregulation of photosynthetic machinery and energy conservation genes, and upregulation of stress-responsive elements such as expansins, flavodoxins, and osmoprotectants. Lipidomic profiling showed accumulation of triacylglycerols (TAGs) and degradation of galactosyl lipids, consistent with a shift toward lipid biosynthesis to mitigate redox imbalance. Depletion of key polar metabolites and branched-chain amino acids suggested a rerouting of central carbon metabolism under stress. DAP-seq identified key transcription factors, including LHY1 and SPL12, that target central metabolic enzymes involved in redox balancing, such as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and malate dehydrogenase (MDH). These findings establish a regulatory-metabolic framework linking redox stress to lipid accumulation and reveal potential engineering targets to enhance salt tolerance. Overall, the multi-omic analysis supports the “overflow” hypothesis, where impaired photosynthesis results in excess reducing equivalents being diverted into TAG synthesis and highlights transcriptional regulators as candidates for improving algal robustness in brackish environments.

09 BIOMASS FUELS↗

IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata

Viruses are widely recognized as critical members of all microbiomes. Metagenomics enables large-scale exploration of the global virosphere, progressively revealing the extensive genomic diversity of viruses on Earth and highlighting the myriad of ways by which viruses impact biological processes. IMG/VR provides access to the largest collection of viral sequences obtained from (meta)genomes, along with functional annotation and rich metadata. A web interface enables users to efficiently browse and search viruses based on genome features and/or sequence similarity. Here, for this work, we present the fourth version of IMG/VR, composed of >15 million virus genomes and genome fragments, a ≈6-fold increase in size compared to the previous version. These clustered into 8.7 million viral operational taxonomic units, including 231 408 with at least one high-quality representative. Viral sequences in IMG/VR are now systematically identified from genomes, metagenomes, and metatranscriptomes using a new detection approach (geNomad), and IMG standard annotation are complemented with genome quality estimation using CheckV, taxonomic classification reflecting the latest taxonomic standards, and microbial host taxonomy prediction. IMG/VR v4 is available at https://img.jgi.doe.gov/vr, and the underlying data are available to download at https://genome.jgi.doe.gov/portal/IMG_VR.

59 BASIC BIOLOGICAL SCIENCES↗

The architecture of resilience: a genome assembly of Myrothamnus flabellifolia sheds light on desiccation tolerance and sex determination

Myrothamnus flabellifolia is a dioecious resurrection plant endemic to southern Africa that has become an important model for understanding desiccation tolerance. Despite its ecological and medicinal significance, genomic and transcriptomic resources for the species are limited. We generated a chromosome-level, haplotype-resolved reference genome assembly and annotation for M. flabellifolia and conducted transcriptomic profiling across a natural dehydration–rehydration time course in the field. Genome architecture and sex determination were characterized, and co-expression network and cis-regulatory element (CRE) enrichment analyses were used to investigate dynamic responses to desiccation. The 1.28-Gb genome exhibits unusually consistent chromatin architecture with unique chromosome organization across highly divergent haplotypes. We identified an XY sexual system with a small sex-determining region on Chromosome 8. Transcriptomic responses varied with dehydration severity, pointing to early suppression of growth, progressive activation of protective mechanisms, and subsequent return to homeostasis upon rehydration. Late embryogenesis abundant and early light-induced protein transcripts were dynamically regulated and showed enrichment of abscisic acid and stress-responsive CREs pointing toward conserved responses. Together, this study provides foundational resources for understanding the genomic architecture and reproductive biology of M. flabellifolia and offers new insights into the mechanisms of desiccation tolerance.

chromosome structure↗

Chromosome‐level Thlaspi arvense genome provides new tools for translational research and for a newly domesticated cash cover crop of the cooler climates

Summary Thlaspi arvense (field pennycress) is being domesticated as a winter annual oilseed crop capable of improving ecosystems and intensifying agricultural productivity without increasing land use. It is a selfing diploid with a short life cycle and is amenable to genetic manipulations, making it an accessible field‐based model species for genetics and epigenetics. The availability of a high‐quality reference genome is vital for understanding pennycress physiology and for clarifying its evolutionary history within the Brassicaceae. Here, we present a chromosome‐level genome assembly of var. MN106‐Ref with improved gene annotation and use it to investigate gene structure differences between two accessions (MN108 and Spring32‐10) that are highly amenable to genetic transformation. We describe non‐coding RNAs, pseudogenes and transposable elements, and highlight tissue‐specific expression and methylation patterns. Resequencing of forty wild accessions provided insights into genome‐wide genetic variation, and QTL regions were identified for a seedling colour phenotype. Altogether, these data will serve as a tool for pennycress improvement in general and for translational research across the Brassicaceae.

59 BASIC BIOLOGICAL SCIENCES↗

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Rapid Annotation of Photosynthetic Systems (RAPS): automated algorithm to generate genome-scale metabolic networks from algal genomes

Algae have mounting potential to serve as platform strains for engineering efforts, but tools for these species lag behind traditional model systems like E. coli and yeast. Metabolic models are one such tool that would enhance our ability to rationally engineer algae but current automated algorithms to build models do not perform well for algal genomes. We introduce a novel software pipeline, Rapid Annotation of Photosynthetic Systems (RAPS), which leverages manual curation efforts of published models to create high quality first draft metabolic networks for new species. We compared models produced by our pipeline to published models and found that these models are able to capture more genes than the published models and are more predictive of experimentally determined growth rate. We used RAPS to automate the creation of 8 first draft metabolic models of new species to enable the further study of metabolism in these species.

09 BIOMASS FUELS↗

A Chromosome-Scale Genome Assembly of a Helicoverpa zea Strain Resistant to Bacillus thuringiensis Cry1Ac Insecticidal Protein

Helicoverpa zea (Lepidoptera: Noctuidae) is an insect pest of major cultivated crops in North and South America. The species has adapted to different host plants and developed resistance to several insecticidal agents, including Bacillus thuringiensis (Bt) insecticidal proteins in transgenic cotton and maize. Helicoverpa zea populations persist year-round in tropical and subtropical regions, but seasonal migrations into temperate zones increase the geographic range of associated crop damage. To better understand the genetic basis of these physiological and ecological characteristics, we generated a high-quality chromosome-level assembly for a single H. zea male from Bt-resistant strain, HzStark_Cry1AcR. Hi-C data were used to scaffold an initial 375.2 Mb contig assembly into 30 autosomes and the Z sex chromosome (scaffold N50 = 12.8 Mb and L50 = 14). The scaffolded assembly was error-corrected with a novel pipeline, polishCLR. The mitochondrial genome was assembled through an improved pipeline and annotated. Assessment of this genome assembly indicated 98.8% of the Lepidopteran Benchmark Universal Single-Copy Ortholog set were complete (98.5% as complete single copy). Repetitive elements comprised approximately 29.5% of the assembly with the plurality (11.2%) classified as retroelements. This chromosome-scale reference assembly for H. zea, ilHelZeax1.1, will facilitate future research to evaluate and enhance sustainable crop production practices.

59 BASIC BIOLOGICAL SCIENCES↗

Compilation and utilization of a sorghum transcriptome compendium for gene regulatory network analysis and crop trait engineering

Sorghum bicolor (Sorghum) is a drought and heat tolerant C4 grass crop used to produce grain, forage, biofuels, and other bioproducts. Genetic improvement of sorghum hybrid crops is aided by a large and diverse germplasm, sorghum's diploid inbreeding genetics, and a relatively small genome that has facilitated genomic research. Over the past 20 years, the sorghum research community characterized the cytogenetic and recombinant landscapes of sorghum's 10 chromosomes, sequenced and annotated the sorghum genome, and used that information to identify genes/alleles that modulate flowering time, plant height, seed shattering, and other important traits. More recently, >1000 RNA-seq transcriptome profiles were collected from 15 sorghum genotypes to help understand the genetic basis of variation in growth and development of sorghum stems, tillers, roots, and leaves, and the regulation of biosynthetic pathways that produce epicuticular wax, dhurrin, and RFOs, compounds that contribute to sorghum's resilience. Transcriptome studies were designed to identify differentially expressed genes that are co-expressed during development or in response to a treatment to enable construction of gene regulatory networks. Co-expression and network analysis identified transcription factors and their cognate binding sites in target gene promoters and signaling pathways that modulate gene regulatory networks providing gene editing targets for further trait optimization. RNA-seq data from >20 experiments targeting sorghum organs, tissues, cell types, developmental stages, and responses to environmental conditions (i.e., diel, day-length, shading, water-deficit, temperature) has been compiled in a sorghum transcriptome compendium. The goal of this resource paper is to describe compendium content, accessibility, and a compendium data analysis pipeline and to illustrate the types of information that can be derived from the compendium with a focus on the elucidation of gene regulatory networks useful for guiding the improvement of sorghum traits through gene editing.

RNA-seq↗

Draft genome of Rosenbergiella nectarea strain 8N4 T provides insights into the potential role of this species in its plant host

Background: Rosenbergiella nectarea strain 8N4 T , the type species of the genus Rosenbergiella, was isolated from Amygdalus communis (almond) floral nectar. Other strains of this species were isolated from the floral nectar of Citrus paradisi (grapefruit), Nicotiana glauca (tobacco tree) and from Asphodelus aestivus. R. nectarea strain 8N4 T is a Gram-negative, oxidase-negative, facultatively anaerobic bacterium in the family Enterobacteriaceae. Results: Here we describe features of this organism, together with its genome sequence and annotation. The DNA GC content is 47.38%, the assembly size is 3,294,717 bp, and the total number of genes are 3,346. The genome discloses the possible role that this species may play in the plant. The genome contains both virulence genes, like pectin lyase and hemolysin, that may harm plant cells and genes that are predicted to produce volatile compounds that may impact the visitation rates by nectar consumers, such as pollinators and nectar thieves. Conclusions: The genome of R. nectarea strain 8N4T reveals a mutualistic interaction with the plant host and a possible effect on plant pollination and fitness.

59 BASIC BIOLOGICAL SCIENCES↗