Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

High-quality RNA extraction and the regulation of genes encoding cellulosomes are correlated with growth stage in anaerobic fungi

Anaerobic fungi produce biomass-degrading enzymes and natural products that are important to harness for several biotechnology applications. Although progress has been made in the development of methods for extracting nucleic acids for genomic and transcriptomic sequencing of these fungi, most studies are limited in that they do not sample multiple fungal growth phases in batch culture. In this study, we establish a method to harvest RNA from fungal monocultures and fungal–methanogen co-cultures, and also determine an optimal time frame for high-quality RNA extraction from anaerobic fungi. Based on RNA quality and quantity targets, the optimal time frame in which to harvest anaerobic fungal monocultures and fungal-methanogen co-cultures for RNA extraction was 2-5 days of growth post-inoculation. When grown on cellulose, the fungal strain Anaeromyces robustus cocultivated with the methanogen Methanobacterium bryantii upregulated genes encoding fungal carbohydrate-active enzymes and other cellulosome components relative to fungal monocultures during this time frame, but expression patterns changed at 24-hour intervals throughout the fungal growth phase. These results demonstrate the importance of establishing methods to extract high-quality RNA from anaerobic fungi at multiple time points during batch cultivation.

Brown, Jennifer L.↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

Application of prophage sequence analysis to investigate a disease outbreak involving Salmonella Adjame, a rare serovar and implications for the population structure

Introduction Outbreak investigation of foodborne salmonellosis is hindered when the food source is contaminated by multiple strains of Salmonella , creating difficulties matching an incriminated organism recovered from patients with the specific strain in the suspect food. An outbreak of the rare Salmonella Adjame was caused by multiple strains of the organism as revealed by single-nucleotide polymorphism (SNP) variation. The use of highly discriminatory prophage analysis to characterize strains of Salmonella should enable a more precise strain characterization and aid the investigation of foodborne salmonellosis. Methods We have carried out genomic analysis of S. Adjame strains recovered during the course of a recent outbreak and compared them with other strains of the organism ( n = 38 strains), using SNPs to evaluate strain differences present in the core genome, and prophage sequence typing (PST) to evaluate the accessory genome. Phylogenetic analyses were performed using both total prophage content and conserved prophages. Results The PST analysis of the S. Adjame isolates showed a high degree of strain heterogeneity. We observed small clusters made up of 2-6 isolates ( n = 27) and singletons ( n = 11) in stark contrast with the three clusters observed by SNP analysis. In total, we detected 24 prophages of which only four were highly prevalent, namely: Entero_p88 (36/38 strains), Salmon_SEN34 (35/38 strains), Burkho_phiE255 (33/38 strains) and Edward_GF (28/38 strains). Despite the marked strain diversity seen with prophage analysis, the distribution of the four most common prophages matched the clustering observed using core genome. Discussion Mutations in the core and accessory genomes of S. Adjame have shed light on the evolutionary relationships among the Adjame strains and demonstrated a convergence of the variations observed in both fractions of the genome. We conclude that core and accessory genomes analyses should be adopted in foodborne bacteria outbreak investigations to provide a more accurate strain description and facilitate reliable matching of isolates from patients and incriminated food sources. The outcomes should translate to a better understanding of the microbial population structure and an 46 improved source attribution in foodborne illnesses.

Gao, Ruimin↗

Data-Driven Whole-Genome Clustering to Detect Geospatial, Temporal, and Functional Trends in SARS-CoV-2 Evolution

Current methods for defining SARS-CoV-2 lineages ignore the vast majority of the SARS-CoV-2 genome. We develop and apply an exhaustive vector comparison method that directly compares all known SARS-CoV-2 genome sequences to produce novel lineage classifications. We utilize data-driven models that (i) accurately capture the complex interactions across the set of all known SARS-CoV-2 genomes, (ii) scale to leadership-class computing systems, and (iii) enable tracking how such strains evolve geospatially over time. We show that during the height of the original Omicron surge, countries across Europe, Asia, and the Americas had a spatially asynchronous distribution of Omicron sub-strains. Moreover, neighboring countries were often dominated by either different clusters of the same variant or different variants altogether throughout the pandemic. Analyses of this kind may suggest a different pattern of epidemiological risk than was understood from conventional data, as well as produce actionable insights and transform our ability to prepare for and respond to current and future biological threats.

Jacobson, Daniel↗

Detecting operons in bacterial genomes via visual representation learning

Contiguous genes in prokaryotes are often arranged into operons. Detecting operons plays a critical role in inferring gene functionality and regulatory networks. Human experts annotate operons by visually inspecting gene neighborhoods across pileups of related genomes. These visual representations capture the inter-genic distance, strand direction, gene size, functional relatedness, and gene neighborhood conservation, which are the most prominent operon features mentioned in the literature. By studying these features, an expert can then decide whether a genomic region is part of an operon. We propose a deep learning based method named Operon Hunter that uses visual representations of genomic fragments to make operon predictions. Using transfer learning and data augmentation techniques facilitates leveraging the powerful neural networks trained on image datasets by re-training them on a more limited dataset of extensively validated operons. Our method outperforms the previously reported state-of-the-art tools, especially when it comes to predicting full operons and their boundaries accurately. Furthermore, our approach makes it possible to visually identify the features influencing the network’s decisions to be subsequently cross-checked by human experts.

59 BASIC BIOLOGICAL SCIENCES↗

BSMV-mediated genome editing exhibits host-specific heritability: germline transmission in barley and somatic edits in Nicotiana benthamiana

Plant RNA virus–mediated guide RNA (gRNA) delivery represents a transformative advance in genome editing technologies. Unlike conventional transformation methods that rely on labor-intensive tissue culture and regeneration for each individual gRNA delivery, viral vectors can rapidly and systemically transmit gRNAs into pre-established Cas-expressing plants, providing an accelerated route for functional genomics and trait discovery directly in planta . However, key design parameters, including subgenomic promoter choice, transcript architecture, and their effects on viral fitness and editing outcomes, remain to be elucidated for most viral platforms. We developed five Barley stripe mosaic virus (BSMV) vectors, each with distinct subgenomic promoter elements to drive single gRNA expression. These were initially evaluated in Cas9-expressing transgenic Nicotiana benthamiana plants targeting the Phytoene desaturase ( PDS ) gene to compare their editing efficiencies. Single gRNAs expressed under the duplicated γb subgenomic promoter or when fused directly to the γb genome achieved the highest mutation frequencies (up to 90% at 60 days post-inoculation), whereas β1- and β2-driven sgRNAs produced delayed and reduced editing. Thus, promoter selection critically determines gRNA accumulation and the efficacy of BSMV-mediated genome editing. The top-performing design was then applied to Cas9-expressing barley ( Hordeum vulgare ) targeting HvCMF7 (conferring green-white variegation) and HvGW2.1 (impacts grain width and weight). BSMV spread systemically throughout barley, inducing somatic and heritable mutations at frequencies up to 100%, with virus-free edited progeny. In contrast, despite robust somatic editing in N. benthamiana, no heritable mutations were detected indicating species-dependent limitations in germline transmission. Our systematic comparison of subgenomic promoter architectures establishes clear design principles for optimizing viral vector–mediated delivery. Promoter choice and transcript structure critically shape editing efficiency and viral stability. The host-specific boundary for germline editing, defined by efficient heritable editing in barley but not N. benthamiana , highlights where BSMV offers advantages and where alternative vectors or hybrid strategies are required, guiding rational platform selection for diverse crop species and applications. Collectively, these findings establish BSMV as a promising next-generation vector for rapid, tissue culture–free, and transformation-independent genome editing in cereals and other recalcitrant monocots.

barley↗

Accuracy-Based Annotation Quality Score (ABAQS) v1.0

Assessing genome annotation quality is crucial for downstream analyses, but current methods are inadequate for eukaryotes. We present Accuracy-Based Annotation Quality Score (ABAQS), a novel, minimal-data-driven method that comprehensively assesses annotation quality. ABAQS evaluates multiple factors, including genome completeness, gene model validity, and protein profile accuracy, outperforming other metrics like BUSCO and PSAURON. We applied ABAQS to over 2500 eukaryotic genomes and showed its robustness and effectiveness in evaluating genome annotation quality, making it a valuable tool for researchers working with genomic data. ABAQS reveals significant variation in annotation quality and highlights the importance of filtering in improving annotation quality and accuracy.

Haridas, Sajeet [Lawrence Berkeley National Labora↗

Addressing the pervasive scarcity of structural annotation in eukaryotic algae

Abstract Despite a continuous increase in algal genome sequencing, structural annotations of most algal genome assemblies remain unavailable. This pervasive scarcity of genome annotation has restricted rigorous investigation of these genomic resources and may have precipitated misleading biological interpretations. However, the annotation process for eukaryotic algal species is often challenging as genomic resources and transcriptomic evidence are not always available. To address this challenge, we benchmark the cutting-edge gene prediction methods that can be generalized for a broad range of non-model eukaryotes. Using the most accurate methods selected based on high-quality algal genomes, we predict structural annotations for 135 unannotated algal genomes. Using previously available genomic data pooled together with new data obtained in this study, we identified the core orthologous genes and the multi-gene phylogeny of eukaryotic algae, including of previously unexplored algal species. This study not only provides a benchmark for the use of structural annotation methods on a variety of non-model eukaryotes, but also compensates for missing data in the current spectrum of algal genomic resources. These results bring us one step closer to the full potential of eukaryotic algal genomics.

59 BASIC BIOLOGICAL SCIENCES↗

Biases in genome reconstruction from metagenomic data

Background Advances in sequencing, assembly, and assortment of contigs into species-specific bins has enabled the reconstruction of genomes from metagenomic data (MAGs). Though a powerful technique, it is difficult to determine whether assembly and binning techniques are accurate when applied to environmental metagenomes due to a lack of complete reference genome sequences against which to check the resulting MAGs. Methods We compared MAGs derived from an enrichment culture containing ~20 organisms to complete genome sequences of 10 organisms isolated from the enrichment culture. Factors commonly considered in binning software—nucleotide composition and sequence repetitiveness—were calculated for both the correctly binned and not-binned regions. This direct comparison revealed biases in sequence characteristics and gene content in the not-binned regions. Additionally, the composition of three public data sets representing MAGs reconstructed from the Tara Oceans metagenomic data was compared to a set of representative genomes available through NCBI RefSeq to verify that the biases identified were observable in more complex data sets and using three contemporary binning software packages. Results Repeat sequences were frequently not binned in the genome reconstruction processes, as were sequence regions with variant nucleotide composition. Genes encoded on the not-binned regions were strongly biased towards ribosomal RNAs, transfer RNAs, mobile element functions and genes of unknown function. Our results support genome reconstruction as a robust process and suggest that reconstructions determined to be >90% complete are likely to effectively represent organismal function; however, population-level genotypic heterogeneity in natural populations, such as uneven distribution of plasmids, can lead to incorrect inferences.

54 ENVIRONMENTAL SCIENCES↗

A Targeted Sequencing Assay for Serotyping Escherichia coli Using AgriSeq Technology

The gold standard method for serotyping Escherichia coli has relied on antisera-based typing of the O- and H-antigens, which is labor intensive and often unreliable. In the post-genomic era, sequence-based assays are potentially faster to provide results, could combine O-serogrouping and H-typing in a single test, and could simultaneously screen for the presence of other genetic markers of interest such as virulence factors. Whole genome sequencing is one approach; however, this method has limited multiplexing capabilities, and only a small fraction of the sequence is informative for subtyping or identifying virulence potential. A targeted, sequence-based assay and accompanying software for data analysis would be a great improvement over the currently available methods for serotyping. The purpose of this study was to develop a high-throughput, molecular method for serotyping E. coli by sequencing the genes that are required for production of O- and H-antigens, as well as to develop software for data analysis and serotype identification. To expand the utility of the assay, targets for the virulence factors, Shiga toxins ( stx 1 , and stx 2 ) and intimin ( eae ) were included. To validate the assay, genomic DNA was extracted from O-serogroup and H-type standard strains and from Shiga toxin-producing E. coli , the targeted regions were amplified, and then sequencing libraries were prepared from the amplified products followed by sequencing of the libraries on the Ion S5™ sequencer. The resulting sequence files were analyzed via the SeroType Caller™ software for identification of O-serogroup, H-type, and presence of stx 1 , stx 2 , and eae . We successfully identified 169 O-serogroups and 41 H-types. The assay also routinely detected the presence of stx 1a,c,d (3 of 3 strains), stx 2c−e,g (8 of 8 strains), stx 2f (1 strain), and eae (6 of 6 strains). Taken together, the high-throughput, sequence-based method presented here is a reliable alternative to antisera-based serotyping methods for E. coli .

Elder, Jacob R.↗

Synthetic hybrids of six yeast species

Abstract Allopolyploidy generates diversity by increasing the number of copies and sources of chromosomes. Many of the best-known evolutionary radiations, crops, and industrial organisms are ancient or recent allopolyploids. Allopolyploidy promotes differentiation and facilitates adaptation to new environments, but the tools to test its limits are lacking. Here we develop an iterative method of Hybrid Production (iHyPr) to combine the genomes of multiple budding yeast species, generating Saccharomyces allopolyploids of at least six species. When making synthetic hybrids, chromosomal instability and cell size increase dramatically as additional copies of the genome are added. The six-species hybrids initially grow slowly, but they rapidly regain fitness and adapt, even as they retain traits from multiple species. These new synthetic yeast hybrids and the iHyPr method have potential applications for the study of polyploidy, genome stability, chromosome segregation, and bioenergy.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

To have value, comparisons of high-throughput phenotyping methods need statistical tests of bias and variance

The gap between genomics and phenomics is narrowing. The rate at which it is narrowing, however, is being slowed by improper statistical comparison of methods. Quantification using Pearson’s correlation coefficient ( r ) is commonly used to assess method quality, but it is an often misleading statistic for this purpose as it is unable to provide information about the relative quality of two methods. Using r can both erroneously discount methods that are inherently more precise and validate methods that are less accurate. These errors occur because of logical flaws inherent in the use of r when comparing methods, not as a problem of limited sample size or the unavoidable possibility of a type I error. A popular alternative to using r is to measure the limits of agreement (LOA). However both r and LOA fail to identify which instrument is more or less variable than the other and can lead to incorrect conclusions about method quality. An alternative approach, comparing variances of methods, requires repeated measurements of the same subject, but avoids incorrect conclusions. Variance comparison is arguably the most important component of method validation and, thus, when repeated measurements are possible, variance comparison provides considerable value to these studies. Statistical tests to compare variances presented here are well established, easy to interpret and ubiquitously available. The widespread use of r has potentially led to numerous incorrect conclusions about method quality, hampering development, and the approach described here would be useful to advance high throughput phenotyping methods but can also extend into any branch of science. The adoption of the statistical techniques outlined in this paper will help speed the adoption of new high throughput phenotyping techniques by indicating when one should reject a new method, outright replace an old method or conditionally use a new method.

59 BASIC BIOLOGICAL SCIENCES↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing↗

Methods for immortalization of epithelial cells

Methods for inducing non-clonal immortalization of normal epithelial cells by directly targeting the two main senescence barriers encountered by cultured epithelial cells. In finite lifespan pre-stasis human mammary epithelial cells (HMEC), the stress-associated stasis barrier was bypassed, and in post-stasis HMEC, the replicative senescence barrier, a consequence of critically shortened telomeres, was bypassed. Early passage non-clonal immortalized lines exhibited normal karyotypes. Methods of efficient HMEC immortalization, in the absence of “passenger” genomic errors, should facilitate examination of telomerase regulation and immortalization during human carcinoma progression, methods for screening for toxic and environmental effect on progression, and the development of therapeutics targeting the process of immortalization.

59 BASIC BIOLOGICAL SCIENCES↗