Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

High-Throughput Functional Genomics for Energy Production

Functional genomics remains a foundational field for establishing genotype-phenotype relationships that enable strain engineering. High-throughput (HTP) methods accelerate the Design-Build-Test-Learn cycle that currently drives synthetic biology towards a forward engineering future. Trackable mutagenesis techniques including transposon insertion sequencing and CRISPR-Cas-mediated genome editing allow for rapid fitness profiling of a collection, or library, of mutants to discover beneficial mutations. Due to the relative speed of these experiments compared to adaptive evolution experiments, iterative rounds of mutagenesis can be implemented for next-generation metabolic engineering efforts to design complex production and tolerance phenotypes. Further, the expansion of these mutagenesis techniques to novel bacteria are opening up industrial microbes that show promise for establishing a bio-based economy.

59 BASIC BIOLOGICAL SCIENCES↗

ATCCfinder - Download and Search the ATCC Genome Portal

Much strain-specific sequence data exists in research conducted before the deployment of large sequencing repositories, making it challenging to identify and validate the identity of strains used in these studies through bioinformatics and phenotyping. The American Type Culture Collection (ATCC) is an organization that sells a wide variety of microbes with strain-level taxonomy classification and associated sequenced reference genomes. Currently, ATCC does not provide a method for searching for sequence similarity between a query sequence and their database of reference genomes. Here I propose the software ATCCfinder, which utilizes ATCC application interface software (API) to generate query-able databases from ATCC Genome resources.

Koehler, Samuel↗

Ornamental origins and genomic frontiers: a review of big-bracted dogwood research

The big-bracted (Benthamidia) dogwood clade consists of small- to medium-sized deciduous trees within the genus Cornus, known for their showy spring-time floral bract display. Cornus is within the family Cornaceae and order Cornales, and as Cornales is one of the earliest diverging asterids, these taxa have been important for phylogenetic research. Three species within the big-bracted clade, flowering (Cornus florida), kousa (C. kousa), and Pacific (C. nuttallii) dogwoods, are popular ornamental landscape plants in North America, with more than 130 cultivars released. Despite their commercial popularity, numerous research gaps have limited the expansion of fundamental research and dogwood breeding programs. In this present review, we aim to provide a thorough overview of our current understanding of 1) the phylogenetic and biogeographic context, 2) plant biology and major pests and pathogens impacting commercialization, 3) historical commercialization and propagation methods, and 4) genetic and genomic resources and how they have been implemented to understand these species. Research gaps and future directions to advance basic research and breeding of big-bracted ornamental dogwoods are discussed throughout.

Cornus florida↗

Simulating metagenomic stable isotope probing datasets with MetaSIPSim

DNA-stable isotope probing (DNA-SIP) links microorganisms to their in-situ function in diverse environmental samples. Combining DNA-SIP and metagenomics (metagenomic-SIP) allows us to link genomes from complex communities to their specific functions and improves the assembly and binning of these targeted genomes. However, empirical development of metagenomic-SIP methods is hindered by the complexity and cost of these studies. We developed a toolkit, ‘MetaSIPSim,’ to simulate sequencing read libraries for metagenomic-SIP experiments. MetaSIPSim is intended to generate datasets for method development and testing. To this end, we used MetaSIPSim generated data to demonstrate the advantages of metagenomic-SIP over a conventional shotgun metagenomic sequencing experiment. Through simulation we show that metagenomic-SIP improves the assembly and binning of isotopically labeled genomes relative to a conventional metagenomic approach. Improvements were dependent on experimental parameters and on sequencing depth. Community level G + C content impacted the assembly of labeled genomes and subsequent binning, where high community G + C generally reduced the benefits of metagenomic-SIP. Furthermore, when a high proportion of the community is isotopically labeled, the benefits of metagenomic-SIP decline. Finally, the choice of gradient fractions to sequence greatly influences method performance. Metagenomic-SIP is a valuable method for recovering isotopically labeled genomes from complex communities. We show that metagenomic-SIP performance depends on optimization of experimental parameters. MetaSIPSim allows for simulation of metagenomic-SIP datasets which facilitates the optimization and development of metagenomic-SIP experiments and analytical approaches for dealing with these data.

59 BASIC BIOLOGICAL SCIENCES↗

Deconvolute individual genomes from metagenome sequences through short read clustering

Metagenome assembly from short next-generation sequencing data is a challenging process due to its large scale and computational complexity. Clustering short reads by species before assembly offers a unique opportunity for parallel downstream assembly of genomes with individualized optimization. However, current read clustering methods suffer either false negative (under-clustering) or false positive (over-clustering) problems. Here we extended our previous read clustering software, SpaRC, by exploiting statistics derived from multiple samples in a dataset to reduce the under-clustering problem. Using synthetic and real-world datasets we demonstrated that this method has the potential to cluster almost all of the short reads from genomes with sufficient sequencing coverage. The improved read clustering in turn leads to improved downstream genome assembly quality.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Deletion of the cytochrome bc complex from Heliobacterium modesticaldum results in viable but non-phototrophic cells

The heliobacteria, a family of anoxygenic phototrophs, possess the simplest known photosynthetic apparatus. Although they are photoheterotrophs in the light, the heliobacteria can also grow chemotrophically via pyruvate metabolism in the dark. In the heliobacteria, the cytochrome bc complex is responsible for oxidizing menaquinol and reducing cytochrome c 553 in the electron flow cycle used for phototrophy. However, there is no known electron acceptor for the mobile cytochrome c 553 other than the photochemical reaction center. We have, therefore, hypothesized that the cytochrome bc complex is necessary for phototrophy, but unnecessary for chemotrophic growth in the dark. Here, we used a two-step method for CRISPR-based genome editing in Heliobacterium modesticaldum to delete the genes encoding the four major subunits of the cytochrome bc complex. Genotypic analysis verified the deletion of the petCBDA gene cluster encoding the catalytic components of the complex. Spectroscopic studies revealed that re-reduction of cytochrome c 553 after flash-induced photo-oxidation was over 100 times slower in the petCBDA mutant compared to the wild-type. Steady-state levels of oxidized P 800 (the primary donor of the photochemical reaction center) were much higher in the petCBDA mutant at every light level, consistent with a limitation in electron flow to the reaction center. The petCBDA mutant was unable to grow phototrophically on acetate plus CO 2 but could grow chemotrophically on pyruvate as a carbon source similar to the wild-type strain in the dark. The mutants could be complemented by reintroduction of the petCBDA gene cluster on a plasmid expressed from the clostridial eno promoter.

59 BASIC BIOLOGICAL SCIENCES↗

Dissecting the Shared Genetic Architecture of Suicide Attempt, Psychiatric Disorders, and Known Risk Factors

Suicide is a leading cause of death worldwide, and non-fatal suicide attempts, which occur far more frequently, are a major source of disability and social and economic burden. Both have substantial genetic etiology, which is partially shared and partially distinct from that of related psychiatric disorders. Methods: We conducted a genome-wide association study (GWAS) of 29,782 suicide attempt (SA) cases and 519,961 controls in the International Suicide Genetics Consortium. The GWAS of SA was conditioned on psychiatric disorders using GWAS summary statistics via mtCOJO, to remove genetic effects on SA mediated by psychiatric disorders. We investigated the shared and divergent genetic architectures of SA, psychiatric disorders and other known risk factors. Results: Two loci reached genome-wide significance for SA: the major histocompatibility complex and an intergenic locus on chromosome 7, which remained associated with SA after conditioning on psychiatric disorders and replicated in an independent cohort from the Million Veteran Program. This locus has been implicated in risk-taking, smoking, and insomnia. SA showed strong genetic correlation with psychiatric disorders, particularly major depression, and also with smoking, pain, risk-taking, sleep disturbances, lower educational attainment, reproductive traits, lower socioeconomic status and poorer general health. After conditioning on psychiatric disorders, the genetic correlations between SA and psychiatric disorders decreased, whereas those with non-psychiatric traits remained largely unchanged. Conclusions: Our results identify a risk locus that contributes more strongly to SA than other phenotypes and suggest a shared underlying biology between SA and known risk factors that is not mediated by psychiatric disorders.

60 APPLIED LIFE SCIENCES↗

Apomixis in Farmers’ Fields: Overview, Case Studies from Forage Grasses and Considerations for Future Apomictic Crops

Apomixis occurs naturally in several commercially important species from diverse plant families. While in some of these species apomixis is yet to be exploited in breeding schemes aimed at fixing heterosis, genetic progress and cultivar development, in other species apomixis has been integrated at different stages of breeding. Some of the most relevant examples come from the subfamily Panicoideae, the second largest subfamily of the Poaceae, and are the main focus of this review. The subfamily encompasses many tropical and sub-tropical grasses and grains of worldwide economic importance. Apomictic tropical forages are prime examples of how apomixis can be used and exploited in the development of marketable cultivars, which are essential to the meat and milk production industries globally. The main commercial forages used as grass pastures covering millions of hectares in tropical and sub-tropical regions are polyploids exhibiting gametophytic apomixis that belong to the genus Urochloa spp. (brachiariagrasses) and to the species Megathyrsus maximus (guineagrass). Buffel grass (Cenchrus ciliaris) and Paspalum spp. are other important apomictic forages bred and used in these regions. Breeding involves large germplasm collections from the centers of origin of the species, and for most of them, sexually reproducing diploid plants have been found. Chromosomically duplicated plants that maintain sexual reproduction are used in crosses with apomictic genotypes for the development and selection of cultivars to be marketed or used as progenitors in subsequent breeding cycles. The peculiarities of each genus/species breeding programs, the cultivars obtained from these programs, and the impact of use of marker assisted selection in cultivar development are presented. In addition, the test or implementation of new technologies such as high throughput phenotyping, and the use of machine learning methods for trait prediction and genomic selection are positively impacting the selection and speed of development of new polyploid apomictic cultivars. Furthermore, genetic transformation techniques, including genome editing, provide an additional layer for design of tailor-made, customer-oriented cultivars.

Cenchrus↗

Efficient DNA sequence compression with neural networks

Abstract Background The increasing production of genomic data has led to an intensified need for models that can cope efficiently with the lossless compression of DNA sequences. Important applications include long-term storage and compression-based data analysis. In the literature, only a few recent articles propose the use of neural networks for DNA sequence compression. However, they fall short when compared with specific DNA compression tools, such as GeCo2. This limitation is due to the absence of models specifically designed for DNA sequences. In this work, we combine the power of neural networks with specific DNA models. For this purpose, we created GeCo3, a new genomic sequence compressor that uses neural networks for mixing multiple context and substitution-tolerant context models. Findings We benchmark GeCo3 as a reference-free DNA compressor in 5 datasets, including a balanced and comprehensive dataset of DNA sequences, the Y-chromosome and human mitogenome, 2 compilations of archaeal and virus genomes, 4 whole genomes, and 2 collections of FASTQ data of a human virome and ancient DNA. GeCo3 achieves a solid improvement in compression over the previous version (GeCo2) of $2.4\%$, $7.1\%$, $6.1\%$, $5.8\%$, and $6.0\%$, respectively. To test its performance as a reference-based DNA compressor, we benchmark GeCo3 in 4 datasets constituted by the pairwise compression of the chromosomes of the genomes of several primates. GeCo3 improves the compression in $12.4\%$, $11.7\%$, $10.8\%$, and $10.1\%$ over the state of the art. The cost of this compression improvement is some additional computational time (1.7–3 times slower than GeCo2). The RAM use is constant, and the tool scales efficiently, independently of the sequence size. Overall, these values outperform the state of the art. Conclusions GeCo3 is a genomic sequence compressor with a neural network mixing approach that provides additional gains over top specific genomic compressors. The proposed mixing method is portable, requiring only the probabilities of the models as inputs, providing easy adaptation to other data compressors or compression-based data analysis tools. GeCo3 is released under GPLv3 and is available for free download at https://github.com/cobilab/geco3.

Silva, Milton↗

Quantification of Cas9 binding and cleavage across diverse guide sequences maps landscapes of target engagement

The RNA-guided nuclease Cas9 has unlocked powerful methods for perturbing both the genome through targeted DNA cleavage and the regulome through targeted DNA binding, but limited biochemical data have hampered efforts to quantitatively model sequence perturbation of target binding and cleavage across diverse guide sequences. We present scalable, sequencing-based platforms for high-throughput filter binding and cleavage and then perform 62,444 quantitative binding and cleavage assays on 35,047 on- and off-target DNA sequences across 90 Cas9 ribonucleoproteins (RNPs) loaded with distinct guide RNAs. We observe that binding and cleavage efficacy, as well as specificity, vary substantially across RNPs; canonically studied guides often have atypically high specificity; sequence context surrounding the target modulates Cas9 on-rate; and Cas9 RNPs may sequester targets in nonproductive states that contribute to “proofreading” capability. Lastly, we distill our findings into an interpretable biophysical model that predicts changes in binding and cleavage for diverse target sequence perturbations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Simple, Cost-Effective, and Automation-Friendly Direct PCR Approach for Bacterial Community Analysis

Understanding bacterial interactions and assembly in complex microbial communities using 16S rRNA sequencing normally requires a large experimental load. However, the current DNA extraction methods, including cell disruption and genomic DNA purification, are normally biased, costly, time-consuming, labor-intensive, and not amenable to miniaturization by droplets or 1,536-well plates due to the significant DNA loss during the purification step for tiny-volume and low-cell-density samples.

16S rRNA sequencing↗

High-quality RNA extraction and the regulation of genes encoding cellulosomes are correlated with growth stage in anaerobic fungi

Anaerobic fungi produce biomass-degrading enzymes and natural products that are important to harness for several biotechnology applications. Although progress has been made in the development of methods for extracting nucleic acids for genomic and transcriptomic sequencing of these fungi, most studies are limited in that they do not sample multiple fungal growth phases in batch culture. In this study, we establish a method to harvest RNA from fungal monocultures and fungal–methanogen co-cultures, and also determine an optimal time frame for high-quality RNA extraction from anaerobic fungi. Based on RNA quality and quantity targets, the optimal time frame in which to harvest anaerobic fungal monocultures and fungal-methanogen co-cultures for RNA extraction was 2-5 days of growth post-inoculation. When grown on cellulose, the fungal strain Anaeromyces robustus cocultivated with the methanogen Methanobacterium bryantii upregulated genes encoding fungal carbohydrate-active enzymes and other cellulosome components relative to fungal monocultures during this time frame, but expression patterns changed at 24-hour intervals throughout the fungal growth phase. These results demonstrate the importance of establishing methods to extract high-quality RNA from anaerobic fungi at multiple time points during batch cultivation.

Brown, Jennifer L.↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

Application of prophage sequence analysis to investigate a disease outbreak involving Salmonella Adjame, a rare serovar and implications for the population structure

Introduction Outbreak investigation of foodborne salmonellosis is hindered when the food source is contaminated by multiple strains of Salmonella , creating difficulties matching an incriminated organism recovered from patients with the specific strain in the suspect food. An outbreak of the rare Salmonella Adjame was caused by multiple strains of the organism as revealed by single-nucleotide polymorphism (SNP) variation. The use of highly discriminatory prophage analysis to characterize strains of Salmonella should enable a more precise strain characterization and aid the investigation of foodborne salmonellosis. Methods We have carried out genomic analysis of S. Adjame strains recovered during the course of a recent outbreak and compared them with other strains of the organism ( n = 38 strains), using SNPs to evaluate strain differences present in the core genome, and prophage sequence typing (PST) to evaluate the accessory genome. Phylogenetic analyses were performed using both total prophage content and conserved prophages. Results The PST analysis of the S. Adjame isolates showed a high degree of strain heterogeneity. We observed small clusters made up of 2-6 isolates ( n = 27) and singletons ( n = 11) in stark contrast with the three clusters observed by SNP analysis. In total, we detected 24 prophages of which only four were highly prevalent, namely: Entero_p88 (36/38 strains), Salmon_SEN34 (35/38 strains), Burkho_phiE255 (33/38 strains) and Edward_GF (28/38 strains). Despite the marked strain diversity seen with prophage analysis, the distribution of the four most common prophages matched the clustering observed using core genome. Discussion Mutations in the core and accessory genomes of S. Adjame have shed light on the evolutionary relationships among the Adjame strains and demonstrated a convergence of the variations observed in both fractions of the genome. We conclude that core and accessory genomes analyses should be adopted in foodborne bacteria outbreak investigations to provide a more accurate strain description and facilitate reliable matching of isolates from patients and incriminated food sources. The outcomes should translate to a better understanding of the microbial population structure and an 46 improved source attribution in foodborne illnesses.

Gao, Ruimin↗

Data-Driven Whole-Genome Clustering to Detect Geospatial, Temporal, and Functional Trends in SARS-CoV-2 Evolution

Current methods for defining SARS-CoV-2 lineages ignore the vast majority of the SARS-CoV-2 genome. We develop and apply an exhaustive vector comparison method that directly compares all known SARS-CoV-2 genome sequences to produce novel lineage classifications. We utilize data-driven models that (i) accurately capture the complex interactions across the set of all known SARS-CoV-2 genomes, (ii) scale to leadership-class computing systems, and (iii) enable tracking how such strains evolve geospatially over time. We show that during the height of the original Omicron surge, countries across Europe, Asia, and the Americas had a spatially asynchronous distribution of Omicron sub-strains. Moreover, neighboring countries were often dominated by either different clusters of the same variant or different variants altogether throughout the pandemic. Analyses of this kind may suggest a different pattern of epidemiological risk than was understood from conventional data, as well as produce actionable insights and transform our ability to prepare for and respond to current and future biological threats.

Jacobson, Daniel↗

Detecting operons in bacterial genomes via visual representation learning

Contiguous genes in prokaryotes are often arranged into operons. Detecting operons plays a critical role in inferring gene functionality and regulatory networks. Human experts annotate operons by visually inspecting gene neighborhoods across pileups of related genomes. These visual representations capture the inter-genic distance, strand direction, gene size, functional relatedness, and gene neighborhood conservation, which are the most prominent operon features mentioned in the literature. By studying these features, an expert can then decide whether a genomic region is part of an operon. We propose a deep learning based method named Operon Hunter that uses visual representations of genomic fragments to make operon predictions. Using transfer learning and data augmentation techniques facilitates leveraging the powerful neural networks trained on image datasets by re-training them on a more limited dataset of extensively validated operons. Our method outperforms the previously reported state-of-the-art tools, especially when it comes to predicting full operons and their boundaries accurately. Furthermore, our approach makes it possible to visually identify the features influencing the network’s decisions to be subsequently cross-checked by human experts.

59 BASIC BIOLOGICAL SCIENCES↗

BSMV-mediated genome editing exhibits host-specific heritability: germline transmission in barley and somatic edits in Nicotiana benthamiana

Plant RNA virus–mediated guide RNA (gRNA) delivery represents a transformative advance in genome editing technologies. Unlike conventional transformation methods that rely on labor-intensive tissue culture and regeneration for each individual gRNA delivery, viral vectors can rapidly and systemically transmit gRNAs into pre-established Cas-expressing plants, providing an accelerated route for functional genomics and trait discovery directly in planta . However, key design parameters, including subgenomic promoter choice, transcript architecture, and their effects on viral fitness and editing outcomes, remain to be elucidated for most viral platforms. We developed five Barley stripe mosaic virus (BSMV) vectors, each with distinct subgenomic promoter elements to drive single gRNA expression. These were initially evaluated in Cas9-expressing transgenic Nicotiana benthamiana plants targeting the Phytoene desaturase ( PDS ) gene to compare their editing efficiencies. Single gRNAs expressed under the duplicated γb subgenomic promoter or when fused directly to the γb genome achieved the highest mutation frequencies (up to 90% at 60 days post-inoculation), whereas β1- and β2-driven sgRNAs produced delayed and reduced editing. Thus, promoter selection critically determines gRNA accumulation and the efficacy of BSMV-mediated genome editing. The top-performing design was then applied to Cas9-expressing barley ( Hordeum vulgare ) targeting HvCMF7 (conferring green-white variegation) and HvGW2.1 (impacts grain width and weight). BSMV spread systemically throughout barley, inducing somatic and heritable mutations at frequencies up to 100%, with virus-free edited progeny. In contrast, despite robust somatic editing in N. benthamiana, no heritable mutations were detected indicating species-dependent limitations in germline transmission. Our systematic comparison of subgenomic promoter architectures establishes clear design principles for optimizing viral vector–mediated delivery. Promoter choice and transcript structure critically shape editing efficiency and viral stability. The host-specific boundary for germline editing, defined by efficient heritable editing in barley but not N. benthamiana , highlights where BSMV offers advantages and where alternative vectors or hybrid strategies are required, guiding rational platform selection for diverse crop species and applications. Collectively, these findings establish BSMV as a promising next-generation vector for rapid, tissue culture–free, and transformation-independent genome editing in cereals and other recalcitrant monocots.

barley↗