Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Footprints of Worldwide Adaptation in Structured Populations of Drosophila melanogaster Through the Expanded DEST 2.0 Genomic Resource

Abstract Large-scale genomic resources can place genetic variation into an ecologically informed context. To advance our understanding of the population genetics of the fruit fly Drosophila melanogaster, we present an expanded release of the community-generated population genomics resource Drosophila Evolution over Space and Time (DEST 2.0; https://dest.bio/). This release includes 530 high-quality pooled libraries from flies collected across six continents over more than a decade (2009 to 2021), most at multiple time points per year; 211 of these libraries are sequenced and shared here for the first time. We used this enhanced resource to elucidate several aspects of the species' demographic history and identify novel signs of adaptation across spatial and temporal dimensions. For example, we showed that the spatial genetic structure of populations is stable over time, but that drift due to seasonal contractions of population size causes populations to diverge over time. We identified signals of adaptation that vary between continents in genomic regions associated with xenobiotic resistance, consistent with independent adaptation to common pesticides. Moreover, by analyzing samples collected during spring and fall across Europe, we provide new evidence for seasonal adaptation related to loci associated with pathogen response. Furthermore, we have also released an updated version of the DEST genome browser. This is a useful tool for studying spatiotemporal patterns of genetic variation in this classic model system.

Biochemistry & Molecular Biology↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

Packaged delivery of CRISPR–Cas9 ribonucleoproteins accelerates genome editing

Effective genome editing requires a sufficient dose of CRISPR–Cas9 ribonucleoproteins (RNPs) to enter the target cell while minimizing immune responses, off-target editing, and cytotoxicity. Clinical use of Cas9 RNPs currently entails electroporation into cells ex vivo, but no systematic comparison of this method to packaged RNP delivery has been made. Here we compared two delivery strategies, electroporation and enveloped delivery vehicles (EDVs), to investigate the Cas9 dosage requirements for genome editing. Using fluorescence correlation spectroscopy, we determined that >1300 Cas9 RNPs per nucleus are typically required for productive genome editing. EDV-mediated editing was >30-fold more efficient than electroporation, and editing occurs at least 2-fold faster for EDV delivery at comparable total Cas9 RNP doses. We hypothesize that differences in efficacy between these methods result in part from the increased duration of RNP nuclear residence resulting from EDV delivery. Our results directly compare RNP delivery strategies, showing that packaged delivery could dramatically reduce the amount of CRISPR–Cas9 RNPs required for experimental or clinical genome editing.

60 APPLIED LIFE SCIENCES↗

Genomic insights into local adaptation and migration success in reintroduced Coho Salmon of the Wenatchee River basin

ABSTRACT Objective Reintroduction of salmonids into regions where they have been extirpated is a common conservation strategy that is often implemented through natural recolonization, translocation of natural populations, or hatchery-based programs. Locally adapting to specific environmental conditions is critical for long-term population viability, particularly for species like Coho Salmon Oncorhynchus kisutch, which face diverse selective pressures during their migration. This study focused on the mid-Columbia River Coho Salmon reintroduction program managed by Yakama Nation Fisheries, which has successfully reintroduced Coho Salmon into the Wenatchee and Methow River basins, Washington. Notably, these populations have adapted to the longer migration route than those in the founding stock, with selection favoring individuals with an earlier arrival time and that can navigate a 15-km, high-gradient canyon to reach optimal spawning grounds. The objectives of this study were to investigate whether specific genomic regions are under selection for traits associated with return location and timing in Coho Salmon. Methods Low-coverage whole-genome resequencing data were used to screen for genomic regions associated with the phenotypes of interest. Results A weak polygenic signal in female Coho Salmon was found to be associated with return group, with a subset of candidate adaptive regions occurring across eight chromosomes. Conclusions These findings provide insights into the genomic mechanisms underlying local adaptation in reintroduced salmon populations and inform broodstock selection strategies aimed at promoting natural production and long-term population sustainability.

Horn, Rebekah L.↗

Complete genomes of Asgard archaea reveal diverse integrated and mobile genetic elements

Asgard archaea are of great interest as the progenitors of Eukaryotes, but little is known about the mobile genetic elements (MGEs) that may shape their ongoing evolution. Here, we describe MGEs that replicate in Atabeyarchaeia, a wetland Asgard archaea lineage represented by two complete genomes. We used soil depth–resolved population metagenomic data sets to track 18 MGEs for which genome structures were defined and precise chromosome integration sites could be identified for confident host linkage. Additionally, we identified a complete 20.67 kbp circular plasmid and two family-level groups of viruses linked to Atabeyarchaeia, via CRISPR spacer targeting. Closely related 40 kbp viruses possess a hypervariable genomic region encoding combinations of specific genes for small cysteine-rich proteins structurally similar to restriction-homing endonucleases. One 10.9 kbp integrative conjugative element (ICE) integrates genomically into theAtabeyarchaeum deiterrae-1chromosome and has a 2.5 kbp circularizable element integrated within it. The 10.9 kbp ICE encodes an expressed Type IIG restriction-modification system with a sequence specificity matching an active methylation motif identified by Pacific Biosciences (PacBio) high-accuracy long-read (HiFi) metagenomic sequencing. Restriction-modification of Atabeyarchaeia differs from that of another coexisting Asgard archaea, Freyarchaeia, which has few identified MGEs but possesses diverse defense mechanisms, including DISARM and Hachiman, not found in Atabeyarchaeia. Overall, defense systems and methylation mechanisms of Asgard archaea likely modulate their interactions with MGEs, and integration/excision and copy number variation of MGEs in turn enable host genetic versatility.

Biochemistry & Molecular Biology↗

RNAi and genome editing of sugarcane: Progress and prospects

SUMMARY Sugarcane, which provides 80% of global table sugar and 40% of biofuel, presents unique breeding challenges due to its highly polyploid, heterozygous, and frequently aneuploid genome. Significant progress has been made in developing genetic resources, including the recently completed reference genome of the sugarcane cultivar R570 and pan‐genomic resources from sorghum, a closely related diploid species. Biotechnological approaches including RNA interference (RNAi), overexpression of transgenes, and gene editing technologies offer promising avenues for accelerating sugarcane improvement. These methods have successfully targeted genes involved in important traits such as sucrose accumulation, lignin biosynthesis, biomass oil accumulation, and stress response. One of the main transformation methods—biolistic gene transfer or Agrobacterium ‐mediated transformation—coupled with efficient tissue culture protocols, is typically used for implementing these biotechnology approaches. Emerging technologies show promise for overcoming current limitations. The use of morphogenic genes can help address genotype constraints and improve transformation efficiency. Tissue culture‐free technologies, such as spray‐induced gene silencing, virus‐induced gene silencing, or virus‐induced gene editing, offer potential for accelerating functional genomics studies. Additionally, novel approaches including base and prime editing, orthogonal synthetic transcription factors, and synthetic directed evolution present opportunities for enhancing sugarcane traits. These advances collectively aim to improve sugarcane's efficiency as a crop for both sugar and biofuel production. This review aims to discuss the progress made in sugarcane methodologies, with a focus on RNAi and gene editing approaches, how RNAi can be used to inform functional gene targets, and future improvements and applications.

Brant, Eleanor [Agronomy Department, Plant Molecul↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗

genomeocean: a pretrained microbial genome foundational model (genomeoceanLLM) v1.0

We present Genomeocean, a foundational genome language model that represents the microbial genome sequences from complex environmental samples. By training on a large, diverse metagenomic dataset, Genomeocean learns species-specific sequence composition and can generate long, realistic open reading frames (ORFs). Our model employs a Byte-pair-encoding (BPE) tokenization strategy, allowing it to efficiently process large genomic datasets and generate long sequences up to 50kb. We demonstrate that fine-tuning Genomeocean can generate novel gene clusters encoding biosynthetic pathways, showcasing its ability to model both fundamental and complex biological processes. Our work establishes Genomeocean as a powerful tool for understanding microbial genome biology and paves the way for its application in a range of fields, from synthetic biology to microbiome research.

Wang, Zhong [Lawrence Berkeley National Laboratory↗

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES↗

Infection and Genomic Properties of Single- and Double-Stranded DNA Cellulophaga Phages

Bacterial viruses (phages) are abundant and ecologically impactful, but laboratory-based experimental model systems vastly under-represent known phage diversity, particularly for ssDNA phages. Here, we characterize the genomes and infection properties of two unrelated marine flavophages—ssDNA generalist phage phi18:4 (6.5 Kbp) and dsDNA specialist phage phi18:1 (39.2 Kbp)—when infecting the same Cellulophaga baltica strain #18 (Cba18), of the class Flavobacteriia. Phage phi18:4 belongs to a new family of ssDNA phages, has an internal lipid membrane, and its genome encodes primarily structural proteins, as well as a DNA replication protein common to ssDNA phages and a unique lysis protein. Phage phi18:1 is a siphovirus that encodes several virulence genes, despite not having a known temperate lifestyle, a CAZy enzyme likely for regulatory purposes, and four DNA methyltransferases dispersed throughout the genome that suggest both host modulation and phage DNA protection against host restriction. Physiologically, ssDNA phage phi18:4 has a shorter latent period and smaller burst size than dsDNA phage phi18:1, and both phages efficiently infect this host. These results help augment the diversity of characterized environmental phage–host model systems by studying infections of genomically diverse phages (ssDNA vs. dsDNA) on the same host.

Howard-Varona, Cristina (ORCID:0000000241495818)↗

Methylation Pattern Detection in the Genome of Bacillus Pumilus Strain SAFR-032

Bacillus pumilus SAFR-032, an endospore-forming bacterial strain that was isolated from a spacecraft assembly facility (SAFR), was investigated to determine its methylation pattern (methylome) across the genome in comparison to the previously sequenced reference genome. In addition, a version of SAFR-032 that was flown as spores for 18 months on the International Space Station (ISS) was also investigated for possible genomic changes due to long-duration ISS-flight and to determine if methylation patterns may have changed. Both the genomics and methylomics were conducted using a Nanopore MinION sequencing device. In addition to the omics investigation, the two SAFR-032 strains, ISS flown and non-ISS flown, were compared phenotypically in chamber experiments testing individual environmental insults: ionizing radiation, UV exposure, and cold desiccation (i.e. freeze drying). Results from this study inform on Planetary Protection concerns and will reveal potential DNA damage associated with long-term spaceflight and how such damage may influence survivors after being transported to an extraterrestrial environment, such as Mars.

Serda, Bianca M.↗

Three pairs of fungal Trametes strains isolated from distinct geographic origins show conserved genomic features and adaptive response to plant biomass

The genomes of white-rot fungi hold extended repertoires of enzymes active on virtually all the chemical bonds that intertwine lignocellulose polymers, and several Trametes species have been identified as powerful tools for biorefinery or bioremediation. However, only few studies have addressed the intra-species polymorphism one would expect from fungal strains collected in contrasted environments. We compared the genome sequence of pairs of strains collected in different geographic areas, for each of three fungal species. Using an updated list of the predicted functions for fungal ligno- and cellulolytic enzymes (CAZymes), we observed a high conservation of the gene repertoires among the six strains. We compared the adaptative response of the fungi grown on crystalline cellulose, wheat straw, aspen or pine sawdust by transcriptomics and secretomics. The gene regulation profiles were determined by the species and the substrates, rather than the strain. The secretomes did not show marked differences in the sets of secreted CAZymes after 3 day-growth on the substrates. We identified five transcription factor genes and two sesquiterpenoid synthesis genes induced during growth on lignocellulose. Wider studies using larger sets of strains will be necessary to evaluate the genericity of our findings, and to assess the phenotype diversity one could expect from geographic diversity as compared to taxonomic diversity in Trametes fungi.

Drula, E. [French National Research Institute for ↗

Data for "Enhancing Lipid Production in Plant Cells through Automated High-Throughput Genome Engineering and Phenotyping"

Plant bioengineering is a time-consuming and labor-intensive process with no guarantee of achieving desired traits. Here, we present a fast, automated, scalable, high-throughput pipeline for plant bioengineering (FAST-PB) in maize (Zea mays) and Nicotiana benthamiana. FAST-PB enables genome editing and product characterization by integrating automated biofoundry engineering of callus and protoplast cells with single-cell matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). We first demonstrated that FAST-PB could streamline Golden Gate cloning, with the capacity to construct 96 vectors in parallel. Using FAST-PB in protoplasts, we found that PEG2050 increased transfection efficiency by over 45%. For proof-of-concept, we established a reporter-gene-free method for CRISPR editing and phenotyping via mutation of high chlorophyll fluorescence 136. We show that diverse lipids were enhanced up to 6-fold using CRISPR activation of lipid controlling genes. In callus cells, an automated transformation platform was employed to regenerate plants with enhanced lipid traits through introducing multigene cassettes. Lastly, FAST-PB enabled high-throughput single-cell lipid profiling by integrating MALDI-MS with the biofoundry, protoplast, and callus cells, differentiating engineered and unengineered cells using single-cell lipidomics. These innovations massively increase the throughput of synthetic biology, genome editing, and metabolic engineering and change what is possible using single-cell metabolomics in plants.

AI/ML↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii

Tilletia maclaganii is a smut fungal pathogen that causes significant biomass reduction of switchgrass ( Panicum virgatum ) used for animal forage and biofuel production. Here we present the annotated genome of T. maclaganii , strain Tm001-NY21, estimated at 42.79 Mb in size, in 53 assembled contigs and encoding 10,235 predicted genes. This genome will be important for future comparative studies of Ustilaginales across its geographic and host range.

PacBio↗

Unique trajectory of gene family evolution from genomic analysis of nearly all known species in an ancient yeast lineage

Gene gains and losses are a major driver of genome evolution; their precise characterization can provide insights into the origin and diversification of major lineages. Here, we examined gene family evolution of 1154 genomes from nearly all known species in the medically and technologically important yeast subphylum Saccharomycotina. We found that yeast gene family evolution differs from that of plants, animals, and filamentous ascomycetes, and is characterized by smaller overall gene numbers yet larger gene family sizes for a given gene number. Faster-evolving lineages (FELs) in yeasts experienced significantly higher rates of gene losses—commensurate with a narrowing of metabolic niche breadth—but higher speciation rates than their slower-evolving sister lineages (SELs). Gene families most often lost are those involved in mRNA splicing, carbohydrate metabolism, and cell division and are likely associated with intron loss, metabolic breadth, and non-canonical cell cycle processes. Our results highlight the significant role of gene family contractions in the evolution of yeast metabolism, genome function, and speciation, and suggest that gene family evolutionary trajectories have differed markedly across major eukaryotic lineages.

Comparative Genomics↗