Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bioinformatics tool”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Interpreting omics data with pathway enrichment analysis

Pathway enrichment analysis is indispensable for interpreting omics datasets and generating hypotheses. However, the foundations of enrichment analysis remain elusive to many biologists. Here, in this study, we discuss best practices in interpreting different types of omics data using pathway enrichment analysis and highlight the importance of considering intrinsic features of various types of omics data. We further explain major components that influence the outcomes of a pathway enrichment analysis, including defining background sets and choosing reference annotation databases. To improve reproducibility, we describe how to standardize reporting methodological details in publications. This article aims to serve as a primer for biologists to leverage the wealth of omics resources and motivate bioinformatics tool developers to enhance the power of pathway enrichment analysis.

60 APPLIED LIFE SCIENCES↗

PathTracer Comprehensively Identifies Hypoxia-Induced Dormancy Adaptations in Mycobacterium tuberculosis

Mining large-scale data to discover biologically relevant information remains a challenge despite the rapid development of bioinformatics tools. Here, we have developed a new tool, PathTracer, to identify biologically relevant information flows by mining genome-wide protein–protein interaction networks following integration of gene expression data. PathTracer successfully mines interactions between genes and traces the most perturbed paths of perceived activities under the conditions of the study. Here, we further demonstrated the utility of this tool by identifying adaptation mechanisms of hypoxia-induced dormancy in Mycobacterium tuberculosis (Mtb).

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Cytochromes P450 involved in bacterial RiPP biosyntheses

Abstract Ribosomally synthesized and post-translationally modified peptides (RiPPs) are a large class of secondary metabolites that have garnered scientific attention due to their complex scaffolds with potential roles in medicine, agriculture, and chemical ecology. RiPPs derive from the cleavage of ribosomally synthesized proteins and additional modifications, catalyzed by various enzymes to alter the peptide backbone or side chains. Of these enzymes, cytochromes P450 (P450s) are a superfamily of heme-thiolate proteins involved in many metabolic pathways, including RiPP biosyntheses. In this review, we focus our discussion on P450 involved in RiPP pathways and the unique chemical transformations they mediate. Previous studies have revealed a wealth of P450s distributed across all domains of life. While the number of characterized P450s involved in RiPP biosyntheses is relatively small, they catalyze various enzymatic reactions such as C–C or C–N bond formation. Formation of some RiPPs is catalyzed by more than one P450, enabling structural diversity. With the continuous improvement of the bioinformatic tools for RiPP prediction and advancement in synthetic biology techniques, it is expected that further cytochrome P450-mediated RiPP biosynthetic pathways will be discovered. Summary The presence of genes encoding P450s in gene clusters for ribosomally synthesized and post-translationally modified peptides expand structural and functional diversity of these secondary metabolites, and here, we review the current state of this knowledge.

59 BASIC BIOLOGICAL SCIENCES↗

Bioinformatics and 3D Structural Analysis of the Coronavirus Main Protease Active Site Diversity

Coronaviruses (Coronaviridae) such as SARS‐CoV‐2 (severe acute respiratory syndrome coronavirus) and MERS‐CoV (Middle East respiratory syndrome coronavirus) have been the source of recent outbreaks and global health concerns. While vaccines have been essential for controlling the SARS‐CoV‐2 (COVID‐19) pandemic, it is uncertain whether they will be effective against future coronavirus strains. Therefore, identification or design of a broad‐spectrum drug that targets highly conserved regions of the main protease of multiple coronavirus strains is essential in the long term. As part of a virtual summer research experience with the RCSB PDB, bioinformatics tools were employed to predict and construct 3D models of the coronavirus main protease (MPro) using SARS‐CoV‐2 as the template, with a focus on mutational trends and active sites. This study focused on the active sites of MPro, a cysteine protease essential for viral assembly and replication. Sequence alignments and structure modeling of MPro structures has identified conserved regions across multiple coronavirus strains. Inhibition of MPro halts coronavirus replication, making it an ideal drug target, and studies of MPro may foster and accelerate the discovery of high affinity broad‐spectrum drugs.

Wu Wu, Amy↗

The need for an integrated multi-OMICs approach in microbiome science in the food system

Microbiome science as an interdisciplinary research field has evolved rapidly over the past two decades, becoming a popular topic not only in the scientific community and among the general public, but also in the food industry due to the growing demand for microbiome-based technologies that provide added-value solutions. Microbiome research has expanded in the context of food systems, strongly driven by methodological advances in different -omics fields that leverage our understanding of microbial diversity and function. However, managing and integrating different complex -omics layers are still challenging. Within the Coordinated Support Action MicrobiomeSupport (https://www.microbiomesupport.eu/), a project supported by the European Commission, the workshop “Metagenomics, Metaproteomics and Metabolomics: the need for data integration in microbiome research” gathered 70 participants from different microbiome research fields relevant to food systems, to discuss challenges in microbiome research and to promote a switch from microbiome-based descriptive studies to functional studies, elucidating the biology and interactive roles of microbiomes in food systems. A combination of technologies is proposed. This will reduce the biases resulting from each individual technology and result in a more comprehensive view of the biological system as a whole. Although combinations of different datasets are still rare, advanced bioinformatics tools and artificial intelligence approaches can contribute to understanding, prediction, and management of the microbiome, thereby providing the basis for the improvement of food quality and safety.

60 APPLIED LIFE SCIENCES↗

Pressured cytotoxic T cell epitope strength among SARS-CoV-2 variants correlates with COVID-19 severity

Heterogeneity in susceptibility among individuals to COVID-19 has been evident through the pandemic worldwide. Cytotoxic T lymphocyte (CTL) responses generated against pathogens in certain individuals are known to impose selection pressure on the pathogen, thus driving emergence of new variants. In this study, we probe the role played by host genetic heterogeneity in terms of HLA-genotypes in determining differential COVID-19 severity in patients. We use bioinformatic tools for CTL epitope prediction to identify epitopes under immune pressure. Using HLA-genotype data of COVID-19 patients from a local cohort, we observe that the recognition of pressured epitopes from the parent strain Wuhan-Hu-1 correlates with COVID-19 severity. We also identify and rank list HLA-alleles and epitopes that offer protectivity against severe disease in infected individuals. Finally, we shortlist a set of 6 pressured and protective epitopes that represent regions in the viral proteome that are under high immune pressure across SARS-CoV-2 variants. In conclusion, identification of such epitopes, defined by the distribution of HLA-genotypes among members of a population, could potentially aid in prediction of indigenous variants of SARS-CoV-2 and other pathogens.

60 APPLIED LIFE SCIENCES↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

wastewater_virus

This repo contains software used to clean and assemble high-throughput sequencing data containing viruses. The input is raw illumina sequencing reads and the output is a database of high-quality viral genomes. The specific application is to wastewater viral concentrates but it is not restricted to that sample type. The software is composed of Nextflow workflows and a set of custom Python and bash scripts that call publicly available bioinformatics tools to accomplish obvious tasks in data analysis in a high performance computing environment. For detailed information, please see the repo's README file.

Kantor, Rose [Lawrence Livermore National Laborato↗

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AlloSHP: deconvoluting single homeologous polymorphism for phylogenetic analysis of allopolyploids

Background The genomic and evolutionary study of allopolyploid organisms involves multiple copies of homeologous chromosomes, making their assembly, annotation, and phylogenetic analysis challenging. Bioinformatics tools and protocols have been developed to study polyploid genomes, but sometimes require the assembly of their genomes, or at least the genes, limiting their use. Results We have developed AlloSHP, a command-line tool for detecting and extracting single homeologous polymorphisms (SHPs) from the subgenomes of allopolyploid species. This tool integrates three main algorithms, WGA, VCF2ALIGNMENT and VCF2SYNTENY, and allows the detection of SHPs for the study of diploid-polyploid complexes with available diploid progenitor genomes, without assembling and annotating the genomes of the allopolyploids under study. AlloSHP has been validated on three diploid-polyploid plant complexes, Brachypodium, Brassica, and Triticum-Aegilops, and a set of synthetic hybrid yeasts and their progenitors of the genus Saccharomyces. The results and congruent phylogenies obtained from the four datasets demonstrate the potential of AlloSHP for the evolutionary analysis of allopolyploids with a wide range of ploidy and genome sizes. Conclusions AlloSHP combines the strategies of simultaneous mapping against multiple reference genomes and syntenic alignment of these genomes to call SHPs, using as input data a single VCF file and the reference genomes of the known or closest extant diploid progenitor species. This novel approach provides a valuable tool for the evolutionary study of allopolyploid species, both at the interspecific and intraspecific levels, allowing the simultaneous analysis of a large number of accessions and avoiding the complex process of assembling polyploid genomes.

Allopolyploids↗

Maast: genotyping thousands of microbial strains efficiently

Existing single nucleotide polymorphism (SNP) genotyping algorithms do not scale for species with thousands of sequenced strains, nor do they account for conspecific redundancy. Here we present a bioinformatics tool, Maast, which empowers population genetic meta-analysis of microbes at an unrivaled scale. Maast implements a novel algorithm to heuristically identify a minimal set of diverse conspecific genomes, then constructs a reliable SNP panel for each species, and enables rapid and accurate genotyping using a hybrid of whole-genome alignment and k-mer exact matching. We demonstrate Maast’s utility by genotyping thousands of Helicobacter pylori strains and tracking SARS-CoV-2 diversification.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting drug-metagenome interactions: Variation in the microbial β-glucuronidase level in the human gut metagenomes

Characterizing the gut microbiota in terms of their capacity to interfere with drug metabolism is necessary to achieve drug efficacy and safety. Although examples of drug-microbiome interactions are well-documented, little has been reported about a computational pipeline for systematically identifying and characterizing bacterial enzymes that process particular classes of drugs. The goal of our study is to develop a computational approach that compiles drugs whose metabolism may be influenced by a particular class of microbial enzymes and that quantifies the variability in the collective level of those enzymes among individuals. The present paper describes this approach, with microbial β-glucuronidases as an example, which break down drug-glucuronide conjugates and reactivate the drugs or their metabolites. We identified 100 medications that may be metabolized by β-glucuronidases from the gut microbiome. These medications included morphine, estrogen, ibuprofen, midazolam, and their structural analogues. The analysis of metagenomic data available through the Sequence Read Archive (SRA) showed that the level of β-glucuronidase in the gut metagenomes was higher in males than in females, which provides a potential explanation for the sex-based differences in efficacy and toxicity for several drugs, reported in previous studies. Our analysis also showed that infant gut metagenomes at birth and 12 months of age have higher levels of β-glucuronidase than the metagenomes of their mothers and the implication of this observed variability was discussed in the context of breastfeeding as well as infant hyperbilirubinemia. Overall, despite important limitations discussed in this paper, our analysis provided useful insights on the role of the human gut metagenome in the variability in drug response among individuals. Importantly, this approach exploits drug and metagenome data available in public databases as well as open-source cheminformatics and bioinformatics tools to predict drug-metagenome interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic Insights Into the Microbial Iron Cycle of Subseafloor Habitats

Microbial iron cycling influences the flux of major nutrients in the environment (e.g., through the adsorptive capacity of iron oxides) and includes biotically induced iron oxidation and reduction processes. The ecological extent of microbial iron cycling is not well understood, even with increased sequencing efforts, in part due to limitations in gene annotation pipelines and limitations in experimental studies linking phenotype to genotype. This is particularly true for the marine subseafloor, which remains undersampled, but represents the largest contiguous habitat on Earth. To address this limitation, we used FeGenie, a database and bioinformatics tool that identifies microbial iron cycling genes and enables the development of testable hypotheses on the biogeochemical cycling of iron. Herein, we survey the microbial iron cycle in diverse subseafloor habitats, including sediment-buried crustal aquifers, as well as surficial and deep sediments. We inferred the genetic potential for iron redox cycling in 32 of the 46 metagenomes included in our analysis, demonstrating the prevalence of these activities across underexplored subseafloor ecosystems. We show that while some processes (e.g., iron uptake and storage, siderophore transport potential, and iron gene regulation) are near-universal, others (e.g., iron reduction/oxidation, siderophore synthesis, and magnetosome formation) are dependent on local redox and nutrient status. Additionally, we detected niche-specific differences in strategies used for dissimilatory iron reduction, suggesting that geochemical constraints likely play an important role in dictating the dominant mechanisms for iron cycling. Overall, our survey advances the known distribution, magnitude, and potential ecological impact of microbe-mediated iron cycling and utilization in sub-benthic ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

Thermophilic Geobacillus WSUCF1 Secretome for Saccharification of Ammonia Fiber Expansion and Extractive Ammonia Pretreated Corn Stover

A thermophilic Geobacillus bacterial strain, WSUCF1 contains different carbohydrate-active enzymes (CAZymes) capable of hydrolyzing hemicellulose in lignocellulosic biomass. We used proteomic, genomic, and bioinformatic tools, and genomic data to analyze the relative abundance of cellulolytic, hemicellulolytic, and lignin modifying enzymes present in the secretomes. Results showed that CAZyme profiles of secretomes varied based on the substrate type and complexity, composition, and pretreatment conditions. The enzyme activity of secretomes also changed depending on the substrate used. The secretomes were used in combination with commercial and purified enzymes to carry out saccharification of ammonia fiber expansion (AFEX)-pretreated corn stover and extractive ammonia (EA)-pretreated corn stover. When WSUCF1 bacterial secretome produced at different conditions was combined with a small percentage of commercial enzymes, we observed efficient saccharification of EA-CS, and the results were comparable to using a commercial enzyme cocktail (87% glucan and 70% xylan conversion). It also opens the possibility of producing CAZymes in a biorefinery using inexpensive substrates, such as AFEX-pretreated corn stover and Avicel, and eliminates expensive enzyme processing steps that are used in enzyme manufacturing. Implementing in-house enzyme production is expected to significantly reduce the cost of enzymes and biofuel processing cost.

Bhalla, Aditya↗

Genome-wide characterization of the soybean DOMAIN OF UNKNOWN FUNCTION 679 membrane protein gene family highlights their potential involvement in growth and stress response

The DMP (DUF679 membrane proteins) family is a plant-specific gene family that encodes membrane proteins. The DMP family genes are suggested to be involved in various programmed cell death processes and gamete fusion during double fertilization in Arabidopsis. However, their functional relevance in other crops remains unknown. This study identified 14 genes from the DMP family in soybean (Glycine max) and characterized their physiochemical properties, subcellular location, gene structure, and promoter regions using bioinformatics tools. Additionally, their tissue-specific and stress-responsive expressions were analyzed using publicly available transcriptome data. Phylogenetic analysis of 198 DMPs from monocots and dicots revealed six clades, with clade-I encoding senescence-related AtDMP1/2 orthologues and clade-II including pollen-specific AtDMP8/9 orthologues. The largest clade, clade-III, predominantly included monocot DMPs, while monocot- and dicot-specific DMPs were assembled in clade-IV and clade-VI, respectively. Evolutionary analysis suggests that soybean GmDMPs underwent purifying selection during evolution. Using 68 transcriptome datasets, expression profiling revealed expression in diverse tissues and distinct responses to abiotic and biotic stresses. The genes Glyma.09G237500 and Glyma.18G098300 showed pistil-abundant expression by qPCR, suggesting they could be potential targets for female organ-mediated haploid induction. Furthermore, cis-acting regulatory elements primarily related to stress-, hormone-, and light-induced pathways regulate GmDMPs, which is consistent with their divergent expression and suggests involvement in growth and stress responses. Overall, our study provides a comprehensive report on the soybean GmDMP family and a framework for further biological functional analysis of DMP genes in soybean or other crops.

59 BASIC BIOLOGICAL SCIENCES↗

Natural Product Gene Clusters in the Filamentous Nostocales Cyanobacterium HT-58-2

Cyanobacteria are known as rich repositories of natural products. One cyanobacterial-microbial consortium (isolate HT-58-2) is known to produce two fundamentally new classes of natural products: the tetrapyrrole pigments tolyporphins A–R, and the diterpenoid compounds tolypodiol, 6-deoxytolypodiol, and 11-hydroxytolypodiol. The genome (7.85 Mbp) of the Nostocales cyanobacterium HT-58-2 was annotated previously for tetrapyrrole biosynthesis genes, which led to the identification of a putative biosynthetic gene cluster (BGC) for tolyporphins. Here, bioinformatics tools have been employed to annotate the genome more broadly in an effort to identify pathways for the biosynthesis of tolypodiols as well as other natural products. A putative BGC (15 genes) for tolypodiols has been identified. Four BGCs have been identified for the biosynthesis of other natural products. Two BGCs related to nitrogen fixation may be relevant, given the association of nitrogen stress with production of tolyporphins. The results point to the rich biosynthetic capacity of the HT-58-2 cyanobacterium beyond the production of tolyporphins and tolypodiols.

60 APPLIED LIFE SCIENCES↗

Application of transport-based metric for continuous interpolation between cryo-EM density maps

Cryogenic electron microscopy (cryo-EM) has become widely used for the past few years in structural biology, to collect single images of macromolecules "frozen in time". As this technique facilitates the identification of multiple conformational states adopted by the same molecule, a direct product of it is a set of 3D volumes, also called EM maps. To gain more insights on the possible mechanisms that govern transitions between different states, and hence the mode of action of a molecule, we recently introduced a bioinformatic tool that interpolates and generates morphing trajectories joining two given EM maps. This tool is based on recent advances made in optimal transport, that allow efficient evaluation of Wasserstein barycenters of 3D shapes. As the overall performance of the method depends on various key parameters, including the sensitivity of the regularization parameter, we performed various numerical experiments to demonstrate how MorphOT can be applied in different contexts and settings. Finally, we discuss current limitations and further potential connections between other optimal transport theories and the conformational heterogeneity problem inherent with cryo-EM data.

3D shapes↗