Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A robust benchmark for detection of germline large deletions and insertions

New technologies and analysis methods are enabling genomic structural variants (SVs) to be detected with ever-increasing accuracy, resolution and comprehensiveness. To help translate these methods to routine research and clinical practice, we developed a sequence-resolved benchmark set for identification of both false-negative and false-positive germline large insertions and deletions. To create this benchmark for a broadly consented son in a Personal Genome Project trio with broadly available cells and DNA, the Genome in a Bottle Consortium integrated 19 sequence-resolved variant calling methods from diverse technologies. The final benchmark set contains 12,745 isolated, sequence-resolved insertion (7,281) and deletion (5,464) calls ≥50 base pairs (bp). The Tier 1 benchmark regions, for which any extra calls are putative false positives, cover 2.51 Gbp and 5,262 insertions and 4,095 deletions supported by ≥1 diploid assembly. We demonstrate that the benchmark set reliably identifies false negatives and false positives in high-quality SV callsets from short-, linked- and long-read sequencing and optical mapping.

59 BASIC BIOLOGICAL SCIENCES↗

Research in Computational Astrobiology

We report on several projects in the field of computational astrobiology, which is devoted to advancing our understanding of the origin, evolution and distribution of life in the Universe using theoretical and computational tools. Research projects included modifying existing computer simulation codes to use efficient, multiple time step algorithms, statistical methods for analysis of astrophysical data via optimal partitioning methods, electronic structure calculations on water-nuclei acid complexes, incorporation of structural information into genomic sequence analysis methods and calculations of shock-induced formation of polycylic aromatic hydrocarbon compounds.

Chaban, Galina↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

Final Technical Report for DE-SC0022206

This project developed foundational genetic, genomic, and epigenetic tools for anaerobic fungi (Neocallimastigomycota), a group of microorganisms with exceptional natural abilities to deconstruct lignocellulosic biomass. Efficient biomass deconstruction remains a major barrier to economical production of renewable fuels, chemicals, and materials from agricultural and forestry residues. The project sought to enable mechanistic studies and future engineering of anaerobic fungi by improving genomic resources, establishing methods for gene expression, and investigating epigenetic regulation of biomass-degrading pathways. Major accomplishments included generation of the first chromosome-scale genome assemblies for multiple anaerobic fungal species, providing publicly available genomic resources that support both engineering and fundamental biological research. The project established the first reproducible system for heterologous gene expression in anaerobic fungi and identified genomic features and mobile genetic elements that may support future development of stable transformation technologies. In parallel, the project demonstrated direct conversion of untreated lignocellulosic biomass into fuels and specialty chemicals through a fungal-yeast bioprocess and identified anaerobic fungal enzymes with utility for metabolic engineering. The research also revealed that epigenetic regulation plays an important role in controlling fungal gene expression and enzyme production, identifying potential strategies for enhancing biomass degradation. Collectively, this work established anaerobic fungi as a tractable emerging platform for bioenergy and biomanufacturing research, generated valuable public resources, trained the next generation of researchers, and advanced DOE-BER goals related to predictive biology, sustainable bioprocessing, and the circular bioeconomy.

Solomon, Kevin [University of Delaware] (ORCID:000↗

Phylogenomics and genetic analysis of solvent-producing Clostridium species

Abstract The genus Clostridium is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 C. beijerinckii , 57 C. saccharobutylicum , 4 C. saccharoperbutylacetonicum , 5 C. butyricum , 7 C. acetobutylicum , and 3 C. tetanomorphum genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or Escherichia coli for enhanced chemical production. Recently, gene sequences from this collection were used to engineer Clostridium autoethanogenum , a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production, as well as butanol, butanoic acid, hexanol and hexanoic acid production.

59 BASIC BIOLOGICAL SCIENCES↗

Imaging and spatially resolved mass spectrometry applications in nephrology

The application of spatially resolved mass spectrometry (MS) and MS imaging approaches for studying biomolecular processes in the kidney is rapidly growing. These powerful methods, which enable label-free and multiplexed detection of many molecular classes across omics domains (including metabolites, drugs, proteins and protein post-translational modifications), are beginning to reveal new molecular insights related to kidney health and disease. Further, the complexity of the kidney often necessitates multiple scales of analysis for interrogating biofluids, whole organs, functional tissue units, single cells and subcellular compartments. Various MS methods can generate omics data across these spatial domains and facilitate both basic science and pathological assessment of the kidney. Optimal processes related to sample preparation and handling for different MS applications are rapidly evolving. Emerging technology and methods, improvement of spatial resolution, broader molecular characterization, multimodal and multiomics approaches and the use of machine learning and artificial intelligence approaches promise to make these applications even more valuable in the field of nephology. Overall, spatially resolved MS and MS imaging methods have the potential to fill much of the omics gap in systems biology analysis of the kidney and provide functional outputs that cannot be obtained using genomics and transcriptomic methods.

60 APPLIED LIFE SCIENCES↗

The NIH Somatic Cell Genome Editing program

The move from reading to writing the human genome offers new opportunities to improve human health. The United States National Institutes of Health (NIH) Somatic Cell Genome Editing (SCGE) Consortium aims to accelerate the development of safer and more-effective methods to edit the genomes of disease-relevant somatic cells in patients, even in tissues that are difficult to reach. Here we discuss the consortium’s plans to develop and benchmark approaches to induce and measure genome modifications, and to define downstream functional consequences of genome editing within human cells. Central to this effort is a rigorous and innovative approach that requires validation of the technology through third-party testing in small and large animals. New genome editors, delivery technologies and methods for tracking edited cells in vivo, as well as newly developed animal models and human biological systems, will be assembled—along with validated datasets—into an SCGE Toolkit, which will be disseminated widely to the biomedical research community. We visualize this toolkit—and the knowledge generated by its applications—as a means to accelerate the clinical development of new therapies for a wide range of conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Budding yeasts in the subphylum Saccharomycotina Genome sequencing and assembly

Eukaryotic life depends on the functional elements encoded by both the nuclear genome and organellar genomes, such as those contained within the mitochondria. The content, size, and structure of the mitochondrial genome varies across organisms with potentially large implications for phenotypic variance and resulting evolutionary trajectories. Among yeasts in the subphylum Saccharomycotina, extensive differences have been observed in various species relative to the model yeast Saccharomyces cerevisiae, but mitochondrial genome sampling across many groups has been scarce, even as hundreds of nuclear genomes have become available. By extracting mitochondrial reads from existing short-read genome sequence datasets, we have greatly expanded both the number of available genomes and the coverage across sparsely sampled clades. Comparison of 353 yeast mitochondrial genomes revealed that, while size and GC content were fairly consistent across species, those in the genera Metschnikowia and Saccharomyces trended larger, while several species in the order Saccharomycetales exhibited lower GC content. Extreme examples for both size and GC content were scattered throughout the subphylum. All mitochondrial genomes shared a core set of protein-coding genes for Complexes III, IV, and V, but they varied in the presence or absence of mitochondrially-encoded canonical Complex I genes. We traced the loss of Complex I genes to a major event in the ancestor of the orders Saccharomycetales and Saccharomycodales, but we also observed several independent losses in the orders Phaffomycetales, Pichiales, and Dipodascales. In contrast to prior hypotheses based on smaller-scale datasets, comparison of evolutionary rates in protein-coding genes showed no bias towards elevated rates among aerobically fermenting (Crabtree/Warburg-positive) yeasts. Mitochondrial introns were widely distributed, but highly enriched in some groups. The majority of mitochondrial introns were poorly conserved within groups, but several were shared within groups, between groups, and even across taxonomic orders, which is consistent with horizontal gene transfer, likely involving homing endonucleases acting as selfish elements. As the number of available fungal nuclear genomes continues to expand, the methods described here to retrieve mitochondrial genome sequences from these datasets will prove invaluable to ensuring that studies of fungal mitochondrial genomes keep pace with their nuclear counterparts.

diversity↗

CoreCruncher : Fast and Robust Construction of Core Genomes in Large Prokaryotic Data Sets

The core genome represents the set of genes shared by all, or nearly all, strains of a given population or species of prokaryotes. Inferring the core genome is integral to many genomic analyses, however, most methods rely on the comparison of all the pairs of genomes; a step that is becoming increasingly difficult given the massive accumulation of genomic data. Here, we present CoreCruncher; a program that robustly and rapidly constructs core genomes across hundreds or thousands of genomes. CoreCruncher does not compute all pairwise genome comparisons and uses a heuristic based on the distributions of identity scores to classify sequences as orthologs or paralogs/xenologs. Although it is much faster than current methods, our results indicate that our approach is more conservative than other tools and less sensitive to the presence of paralogs and xenologs. CoreCruncher is freely available from: https://github.com/lbobay/CoreCruncher. CoreCruncher is written in Python 3.7 and can also run on Python 2.7 without modification. It requires the python library Numpy and either Usearch or Blast. Certain options require the programs muscle or mafft.

59 BASIC BIOLOGICAL SCIENCES↗

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

Tracking ebolavirus genomic drift with a resequencing microarray

Filoviruses are emerging pathogens that cause acute fever with high fatality rate and present a global public health threat. During the 2013–2016 Ebola virus outbreak, genome sequencing allowed the study of virus evolution, mutations affecting pathogenicity and infectivity, and tracing the viral spread. In 2018, early sequence identification of the Ebolavirus as EBOV in the Democratic Republic of the Congo supported the use of an Ebola virus vaccine. However, field-deployable sequencing methods are needed to enable a rapid public health response. Resequencing microarrays (RMA) are a targeted method to obtain genomic sequence on clinical specimens rapidly, and sensitively, overcoming the need for extensive bioinformatic analysis. This study presents the design and initial evaluation of an ebolavirus resequencing microarray (Ebolavirus-RMA) system for sequencing the major genomic regions of four Ebolaviruses that cause disease in humans. The design of the Ebolavirus-RMA system is described and evaluated by sequencing repository samples of three Ebolaviruses and two EBOV variants. The ability of the system to identify genetic drift in a replicating virus was achieved by sequencing the ebolavirus glycoprotein gene in a recombinant virus cultured under pressure from a neutralizing antibody. Comparison of the Ebolavirus-RMA results to the Genbank database sequence file with the accession number given for the source RNA and Ebolavirus-RMA results compared to Next Generation Sequence results of the same RNA samples showed up to 99% agreement.

59 BASIC BIOLOGICAL SCIENCES↗

A Re-Evaluation of African Swine Fever Genotypes Based on p72 Sequences Reveals the Existence of Only Six Distinct p72 Groups

The African swine fever virus (ASFV) is currently causing a world-wide pandemic of a highly lethal disease in domestic swine and wild boar. Currently, recombinant ASF live-attenuated vaccines based on a genotype II virus strain are commercially available in Vietnam. With 25 reported ASFV genotypes in the literature, it is important to understand the molecular basis and usefulness of ASFV genotyping, as well as the true significance of genotypes in the epidemiology, transmission, evolution, control, and prevention of ASFV. Historically, genotyping of ASFV was used for the epidemiological tracking of the disease and was based on the analysis of small fragments that represent less than 1% of the viral genome. The predominant method for genotyping ASFV relies on the sequencing of a fragment within the gene encoding the structural p72 protein. Genotype assignment has been accomplished through automated phylogenetic trees or by comparing the target sequence to the most closely related genotyped p72 gene. To evaluate its appropriateness for the classification of genotypes by p72, we reanalyzed all available genomic data for ASFV. We conclude that the majority of p72-based genotypes, when initially created, were neither identified under any specific methodological criteria nor correctly compared with the already existing ASFV genotypes. Based on our analysis of the p72 protein sequences, we propose that the current twenty-five genotypes, created exclusively based on the p72 sequence, should be reduced to only six genotypes. To help differentiate between the new and old genotype classification systems, we propose that Arabic numerals (1, 2, 8, 9, 15, and 23) be used instead of the previously used Roman numerals. Furthermore, we discuss the usefulness of genotyping ASFV isolates based only on the p72 gene sequence.

59 BASIC BIOLOGICAL SCIENCES↗

Intensity of sample processing methods impacts wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes

Wastewater SARS-CoV-2 surveillance has been deployed since the beginning of the COVID-19 pandemic to monitor the dynamics in virus burden in local communities. Genomic surveillance of SARS-CoV-2 in wastewater, particularly efforts aimed at whole genome sequencing for variant tracking and identification, are still challenging due to low target concentration, complex microbial and chemical background, and lack of robust nucleic acid recovery experimental procedures. The intrinsic sample limitations are inherent to wastewater and are thus unavoidable. Here, we use a statistical approach that couples correlation analyses to a random forest-based machine learning algorithm to evaluate potentially important factors associated with wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes, with a specific focus on the breadth of genome coverage. We collected 182 composite and grab wastewater samples from the Chicago area between November 2020 to October 2021. Samples were processed using a mixture of processing methods reflecting different homogenization intensities (HA + Zymo beads, HA + glass beads, and Nanotrap), and were sequenced using one of the two library preparation kits (the Illumina COVIDseq kit and the QIAseq DIRECT kit). Technical factors evaluated using statistical and machine learning approaches include sample types, certain sample intrinsic features, and processing and sequencing methods. The results suggested that sample processing methods could be a predominant factor affecting sequencing outcomes, and library preparation kits was considered a minor factor. Finally, a synthetic SARS-CoV-2 RNA spike-in experiment was performed to validate the impact from processing methods and suggested that the intensity of the processing methods could lead to different RNA fragmentation

60 APPLIED LIFE SCIENCES↗

Use of Fluorescent Protein Reporters for Assessing and Detecting Genome Editing Reagents and Transgene Expression in Plants

Fluorescent protein reporters have been widely used for monitoring the expression of target genes in various engineered organisms. Although a wide range of analytical approaches (e.g., genotyping PCR, digital PCR, DNA sequencing) have been utilized to detect and identify genome editing reagents and transgene expression in genetically modified plants, these methods are usually limited to use in the late stages of plant transformation and can only be used invasively. Here we describe GFP- and eYGFPuv-based strategies and methods for assessing and detecting genome editing reagents and transgene expression in plants, including protoplast transformation, leaf infiltration, and stable transformation. These methods and strategies enable easy, noninvasive screening of genome editing and transgenic events in plants.

Yuan, Guoliang↗

Exploring the roles of microbes in facilitating plant adaptation to climate change

Plants benefit from their close association with soil microbes which assist in their response to abiotic and biotic stressors. Yet much of what we know about plant stress responses is based on studies where the microbial partners were uncontrolled and unknown. Under climate change, the soil microbial community will also be sensitive to and respond to abiotic and biotic stressors. Thus, facilitating plant adaptation to climate change will require a systems-based approach that accounts for the multi-dimensional nature of plant–microbe–environment interactions. In this perspective, we highlight some of the key factors influencing plant–microbe interactions under stress as well as new tools to facilitate the controlled study of their molecular complexity, such as fabricated ecosystems and synthetic communities. When paired with genomic and biochemical methods, these tools provide researchers with more precision, reproducibility, and manipulability for exploring plant–microbe–environment interactions under a changing climate.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular and Epidemiological Investigation of Fluconazole-resistant Candida parapsilosis —Georgia, United States, 2021

Abstract Background Reports of fluconazole-resistant Candida parapsilosis bloodstream infections are increasing. We describe a cluster of fluconazole-resistant C parapsilosis bloodstream infections identified in 2021 on routine surveillance by the Georgia Emerging Infections Program in conjunction with the Centers for Disease Control and Prevention. Methods Whole-genome sequencing was used to analyze C parapsilosis bloodstream infections isolates. Epidemiological data were obtained from medical records. A social network analysis was conducted using Georgia Hospital Discharge Data. Results Twenty fluconazole-resistant isolates were identified in 2021, representing the largest proportion (34%) of fluconazole-resistant C parapsilosis bloodstream infections identified in Georgia since surveillance began in 2008. All resistant isolates were closely genetically related and contained the Y132F mutation in the ERG11 gene. Patients with fluconazole-resistant isolates were more likely to have resided at long-term acute care hospitals compared with patients with susceptible isolates (P = .01). There was a trend toward increased mechanical ventilation and prior azole use in patients with fluconazole-resistant isolates. Social network analysis revealed that patients with fluconazole-resistant isolates interfaced with a distinct set of healthcare facilities centered around 2 long-term acute care hospitals compared with patients with susceptible isolates. Conclusions Whole-genome sequencing results showing that fluconazole-resistant C parapsilosis isolates from Georgia surveillance demonstrated low genetic diversity compared with susceptible isolates and their association with a facility network centered around 2 long-term acute care hospitals suggests clonal spread of fluconazole-resistant C parapsilosis. Further studies are needed to better understand the sudden emergence and transmission of fluconazole-resistant C parapsilosis.

Misas, Elizabeth (ORCID:0000000162437716)↗

Meta Biome: a multiscale model integrating agent-based and metabolic networks to reveal spatial regulation in gut mucosal microbial communities

ABSTRACT Mucosal microbial communities (MMCs) are complex ecosystems near the mucosal layers of the gut essential for maintaining health and modulating disease states. Despite advances in high-throughput omics technologies, current methodologies struggle to capture the dynamic metabolic interactions and spatiotemporal variations within MMCs. In this work, we presentMetaBiome, a multiscale model integrating agent-based modeling (ABM), finite volume methods, and constraint-based models to explore the metabolic interactions within these communities. Integrating ABM allows for the detailed representation of individual microbial agents each governed by rules that dictate cell growth, division, and interactions with their surroundings. Through a layered approach—encompassing microenvironmental conditions, agent information, and metabolic pathways—we simulated different communities to showcase the potential of the model. Using ourin-silicoplatform, we explored the dynamics and spatiotemporal patterns of MMCs in the proximal small intestine and the cecum, simulating the physiological conditions of the two gut regions. Our findings revealed how specific microbes adapt their metabolic processes based on substrate availability and local environmental conditions, shedding light on spatial metabolite regulation and informing targeted therapies for localized gut diseases.MetaBiome provides a detailed representation of microbial agents and their interactions, surpassing the limitations of traditional grid-based systems. This work marks a significant advancement in microbial ecology, as it offers new insights into predicting and analyzing microbial communities. IMPORTANCE Our study presents a novel multiscale model that combines agent-based modeling, finite volume methods, and genome-scale metabolic models to simulate the complex dynamics of mucosal microbial communities in the gut. This integrated approach allows us to capture spatial and temporal variations in microbial interactions and metabolism that are difficult to study experimentally. Key findings from our model include the following: (i) prediction of metabolic cross-feeding and spatial organization in multi-species communities, (ii) insights into how oxygen gradients and nutrient availability shape community composition in different gut regions, and (iii) identification of spatiallyregulated metabolic pathways and enzymes inE. coli. We believe this work represents a significant advance in computational modeling of microbial communities and provides new insights into the spatial regulation of gut microbiome metabolism. The multiscale modeling approach we have developed could be broadly applicable for studying other complex microbial ecosystems.

Microbiology↗

Bacteriophage-Induced Lipopolysaccharide Mutations in Escherichia coli Lead to Hypersensitivity to Food Grade Surfactant Sodium Dodecyl Sulfate

Bacteriophages (phages) are considered as one of the most promising antibiotic alternatives in combatting bacterial infectious diseases. However, one concern of employing phage application is the emergence of bacteriophage-insensitive mutants (BIMs). Here, we isolated six BIMs from E. coli B in the presence of phage T4 and characterized them using genomic and phenotypic methods. Of all six BIMs, a six-amino acid deletion in glucosyltransferase WaaG likely conferred phage resistance by deactivating the addition of T4 receptor glucose to the lipopolysaccharide (LPS). This finding was further supported by the impaired phage adsorption to BIMs and glycosyl composition analysis which quantitatively confirmed the absence of glucose in the LPS of BIMs. Since LPSs actively maintain outer membrane (OM) permeability, phage-induced truncations of LPSs destabilized the OM and sensitized BIMs to various substrates, especially to the food-grade surfactant sodium dodecyl sulfate (SDS). This hypersensitivity to SDS was exploited to design a T4–SDS combination which successfully prevented the generation of BIMs and eliminated the inoculated bacteria. Collectively, phage-driven modifications of LPSs immunized BIMs from T4 predation but increased their susceptibilities as a fitness cost. The findings of this study suggest a novel strategy to enhance the effectiveness of phage-based food safety interventions.

60 APPLIED LIFE SCIENCES↗