Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bioinformatics tool”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Genome-wide characterization of the soybean DOMAIN OF UNKNOWN FUNCTION 679 membrane protein gene family highlights their potential involvement in growth and stress response

The DMP (DUF679 membrane proteins) family is a plant-specific gene family that encodes membrane proteins. The DMP family genes are suggested to be involved in various programmed cell death processes and gamete fusion during double fertilization in Arabidopsis. However, their functional relevance in other crops remains unknown. This study identified 14 genes from the DMP family in soybean (Glycine max) and characterized their physiochemical properties, subcellular location, gene structure, and promoter regions using bioinformatics tools. Additionally, their tissue-specific and stress-responsive expressions were analyzed using publicly available transcriptome data. Phylogenetic analysis of 198 DMPs from monocots and dicots revealed six clades, with clade-I encoding senescence-related AtDMP1/2 orthologues and clade-II including pollen-specific AtDMP8/9 orthologues. The largest clade, clade-III, predominantly included monocot DMPs, while monocot- and dicot-specific DMPs were assembled in clade-IV and clade-VI, respectively. Evolutionary analysis suggests that soybean GmDMPs underwent purifying selection during evolution. Using 68 transcriptome datasets, expression profiling revealed expression in diverse tissues and distinct responses to abiotic and biotic stresses. The genes Glyma.09G237500 and Glyma.18G098300 showed pistil-abundant expression by qPCR, suggesting they could be potential targets for female organ-mediated haploid induction. Furthermore, cis-acting regulatory elements primarily related to stress-, hormone-, and light-induced pathways regulate GmDMPs, which is consistent with their divergent expression and suggests involvement in growth and stress responses. Overall, our study provides a comprehensive report on the soybean GmDMP family and a framework for further biological functional analysis of DMP genes in soybean or other crops.

59 BASIC BIOLOGICAL SCIENCES↗

Natural Product Gene Clusters in the Filamentous Nostocales Cyanobacterium HT-58-2

Cyanobacteria are known as rich repositories of natural products. One cyanobacterial-microbial consortium (isolate HT-58-2) is known to produce two fundamentally new classes of natural products: the tetrapyrrole pigments tolyporphins A–R, and the diterpenoid compounds tolypodiol, 6-deoxytolypodiol, and 11-hydroxytolypodiol. The genome (7.85 Mbp) of the Nostocales cyanobacterium HT-58-2 was annotated previously for tetrapyrrole biosynthesis genes, which led to the identification of a putative biosynthetic gene cluster (BGC) for tolyporphins. Here, bioinformatics tools have been employed to annotate the genome more broadly in an effort to identify pathways for the biosynthesis of tolypodiols as well as other natural products. A putative BGC (15 genes) for tolypodiols has been identified. Four BGCs have been identified for the biosynthesis of other natural products. Two BGCs related to nitrogen fixation may be relevant, given the association of nitrogen stress with production of tolyporphins. The results point to the rich biosynthetic capacity of the HT-58-2 cyanobacterium beyond the production of tolyporphins and tolypodiols.

60 APPLIED LIFE SCIENCES↗

Application of transport-based metric for continuous interpolation between cryo-EM density maps

Cryogenic electron microscopy (cryo-EM) has become widely used for the past few years in structural biology, to collect single images of macromolecules "frozen in time". As this technique facilitates the identification of multiple conformational states adopted by the same molecule, a direct product of it is a set of 3D volumes, also called EM maps. To gain more insights on the possible mechanisms that govern transitions between different states, and hence the mode of action of a molecule, we recently introduced a bioinformatic tool that interpolates and generates morphing trajectories joining two given EM maps. This tool is based on recent advances made in optimal transport, that allow efficient evaluation of Wasserstein barycenters of 3D shapes. As the overall performance of the method depends on various key parameters, including the sensitivity of the regularization parameter, we performed various numerical experiments to demonstrate how MorphOT can be applied in different contexts and settings. Finally, we discuss current limitations and further potential connections between other optimal transport theories and the conformational heterogeneity problem inherent with cryo-EM data.

3D shapes↗

Two Strategies for Microbial Production of an Industrial Enzyme-Alpha-Amylase

Extremophiles are microorganisms that thrive in, from an anthropocentric view, extreme environments including hot springs, soda lakes and arctic water. This ability of survival at extreme conditions has rendered extremophiles to be of interest in astrobiology, evolutionary biology as well as in industrial applications. Of particular interest to the biotechnology industry are the biological catalysts of the extremophiles, the extremozymes, whose unique stabilities at extreme conditions make them potential sources of novel enzymes in industrial applications. There are two major approaches to microbial enzyme production. This entails enzyme isolation directly from the natural host or creating a recombinant expression system whereby the targeted enzyme can be overexpressed in a mesophilic host. We are employing both methods in the effort to produce alpha-amylases from a hyperthermophilic archaeon (Thermococcus) isolated from a hydrothermal vent in the Atlantic Ocean, as well as from alkaliphilic bacteria (Bacillus) isolated from a soda lake in Tanzania. Alpha-amylases catalyze the hydrolysis of internal alpha-1,4-glycosidic linkages in starch to produce smaller sugars. Thermostable alpha-amylases are used in the liquefaction of starch for production of fructose and glucose syrups, whereas alpha-amylases stable at high pH have potential as detergent additives. The alpha-amylase encoding gene from Thermococcus was PCR amplified using carefully designed primers and analyzed using bioinformatics tools such as BLAST and Multiple Sequence Alignment for cloning and expression in E.coli. Four strains of Bacillus were grown in alkaline starch-enriched medium of which the culture supernatant was used as enzyme source. Amylolytic activity was detected using the starch-iodine method.

Bernhardsdotter, Eva C. M. J.↗

WHONDRS-GUI: a web application for global survey of surface water metabolites

Background The Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) is a consortium that aims to understand complex hydrologic, biogeochemical, and microbial connections within river corridors experiencing perturbations such as dam operations, floods, and droughts. For one ongoing WHONDRS sampling campaign, surface water metabolite and microbiome samples are collected through a global survey to generate knowledge across diverse river corridors. Metabolomics analysis and a suite of geochemical analyses have been performed for collected samples through the Environmental Molecular Sciences Laboratory (EMSL). The obtained knowledge and data package inform mechanistic and data-driven models to enhance predictions of outcomes of hydrologic perturbations and watershed function, one of the most critical components in model-data integration. To support efforts of the multi-domain integration and make the ever-growing data package more accessible for researchers across the world, a Shiny/R Graphical User Interface (GUI) called WHONDRS-GUI was created. Results The web application can be run on any modern web browser without any programming or operational system requirements, thus providing an open, well-structured, discoverable dataset for WHONDRS. Together with a context-aware dynamic user interface, the WHONDRS-GUI has functionality for searching, compiling, integrating, visualizing and exporting different data types that can easily be used by the community. The web application and data package are available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1484811 , which enables users to simultaneously obtain access to the data and code and to subsequently run the web app locally. The WHONDRS-GUI is also available for online use at Shiny Server ( https://xmlin.shinyapps.io/whondrs/ ).

59 BASIC BIOLOGICAL SCIENCES↗

MultiPhATE2: code for functional annotation and comparison of phage genomes

To address a need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of multiPhATE, a functional annotation code released previously, multiPhATE2 performs gene finding using multiple algorithms, compares the results of the algorithms, performs functional annotation of coding sequences, and incorporates additional search algorithms and databases to extend the search space of the original code. MultiPhATE2 performs gene matching among sets of closely related bacteriophage genomes, and uses multiprocessing to speed computations. MultiPhATE2 can be re-started at multiple points within the workflow to allow the user to examine intermediate results and adjust the subsequent computations accordingly. In addition, multiPhATE2 accommodates custom gene calls and sequence databases, again adding flexibility. MultiPhATE2 was implemented in Python 3.7 and runs as a command-line code under Linux or MAC operating systems. Full documentation is provided as a README file and a Wiki website.

59 BASIC BIOLOGICAL SCIENCES↗

Biological Parts Search Portal (BioParts) v1.0.0

BioParts is a web based search portal for biological parts available in the public domain. It combines the ease and convenience of modern web search engines with the capabilities of bioinformatics search tools such as BLAST. This portal, available at bioparts.org, allows anyone to search for publicly accessible biological part information (e.g., NCBI, iGEM, SynBioHub, Addgene), including parts publicly accessible through ICE Registries. Additionally, the portal offers a REST API that enables third-party applications and tools to access the portal's functionality programmatically. While there are several standalone biological part repositories, there doesn't exist an application that indexes these publicly available parts and enables features such as keyword and BLAST searches along with automatic sequence annotation.

Plahar, Hector↗

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat↗

How Cooperative Engagement Programs Strengthen Sequencing Capabilities for Biosurveillance and Outbreak Response

The threat of emerging and re-emerging infectious diseases continues to be a challenge to public and global health security. Cooperative biological engagement programs act to build partnerships and collaborations between scientists and health professionals to strengthen capabilities in biosurveillance. Biosurveillance is the systematic process of detecting, reporting, and responding to especially dangerous pathogens and pathogens of pandemic potential before they become outbreaks, epidemics, and pandemics. One important tool in biosurveillance is next generation sequencing. Expensive sequencing machines, reagents, and supplies make it difficult for countries to adopt this technology. Cooperative engagement programs help by providing funding for technical assistance to strengthen sequencing capabilities. Through workshops and training, countries are able to learn sequencing and bioinformatics, and implement these tools in their biosurveillance programs. Cooperative programs have an important role in building and sustaining collaborations among institutions and countries. One of the most important pieces in fostering these collaborations is trust. Trust provides the confidence that a successful collaboration will benefit all parties involved. With sequencing, this enables the sharing of pathogen samples and sequences. Obtaining global sequencing data helps to identify unknown etiological agents, track pathogen evolution and infer transmission networks throughout the duration of a pandemic. Having sequencing technology in place for biosurveillance generates the capacity to provide real-time data to understand and respond to pandemics. We highlight the need for these programs to continue to strengthen sequencing in biosurveillance. By working together to strengthen sequencing capabilities, trust can be formed, benefitting global health in the face of biological threats.

60 APPLIED LIFE SCIENCES↗

Appendix Q: Recommendations for Developing Molecular Assays for Microbial Pathogen Detection Using Modern In Silico Approaches

We describe the use of in silico approaches to improve the process of molecular assay development and reduce time and cost by utilizing available databases of whole genome pathogen sequences combined with modern bioinformatics and physical modeling tools. Well-characterized assays are needed for accurately detecting pathogens in environmental and patient samples and also for evaluation of the efficacy of a medical countermeasure that may be administered to patients. The polymerase chain reaction (PCR) remains the gold standard for pathogen detection due to the simplicity of its instrumentation, low cost of reagents, and outstanding limit of detection (LOD), sensitivity, and specificity. However, creation of such PCR assays often involves iterations of design, preliminary testing, and thorough validation with clinical isolates and testing in relevant matrices, which can be time consuming, costly, and result in suboptimal assays. Since formal validation (e.g., for Emergency Use Authorization [EUA] or Food and Drug Administration [FDA] licensure) of an infectious disease assay can be very expensive and can require extensive time of development, having a well-designed assay up front is a critical first step. Yet, many assays described in the literature utilized limited design capabilities and many initially promising assays fail the validation process, resulting in increased costs and timelines for successful product development. While the computational approaches outlined in this document by no means obviate the need for wet lab testing, they can reduce the amount of effort wasted on empirical optimization and iterative redesigns and also guide validation studies. The proposed computational approaches also result in higher performing assays with better sensitivity, specificity, and lower LOD and reduce the possibility of assay failure due to signature erosion. To provide clarity, an extensive glossary of defined terms is provided.

59 BASIC BIOLOGICAL SCIENCES↗

AlphaBeta: computational inference of epimutation rates and spectra from high-throughput DNA methylation data in plants

Stochastic changes in DNA methylation (i.e., spontaneous epimutations) contribute to methylome diversity in plants. Here, we describe AlphaBeta, a computational method for estimating the precise rate of such stochastic events using pedigree-based DNA methylation data as input. We demonstrate how AlphaBeta can be employed to study transgenerationally heritable epimutations in clonal or sexually derived mutation accumulation lines, as well as somatic epimutations in long-lived perennials. Application of our method to published and new data reveals that spontaneous epimutations accumulate neutrally at the genome-wide scale, originate mainly during somatic development and that they can be used as a molecular clock for age-dating trees.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗

Global overview and major challenges of host prediction methods for uncultivated phages

Bacterial communities play critical roles across all of Earth’s biomes, affecting human health and global ecosystem functioning. They do so under strong constraints exerted by viruses, i.e., bacteriophages or “phages”. Phages can reshape bacterial communities’ structure, influence long-term evolution of bacterial populations, and alter host cell metabolism during infection. Metagenomics approaches, i.e., shotgun sequencing of environmental DNA or RNA, recently enabled large-scale exploration of phage genomic diversity, yielding several millions of phage genomes now to be further analyzed and characterized. One major challenge however is the lack of direct host information for these phages. Several methods and tools have been proposed to bioinformatically predict the potential host(s) of uncultivated phages based only on genome sequence information. Here we review these different approaches and highlight their distinct strengths and limitations. We also outline complementary experimental assays which are being proposed to validate and refine these bioinformatic predictions.

59 BASIC BIOLOGICAL SCIENCES↗

Flexible and Adaptive Malware Identification Using Techniques from Biology

The holy grail in cyber analytics is to find new ways to understand the information we already have access to. One way to do that is to characterize the data into reasonable sizes and then leverage any known information to generate new insights. Biologists have been using a similar process for decades. This paper introduces the MLSTONES tool set that was developed by leveraging biology and bioinformatics, high performance computing, and statistical algorithms applied to cyber data and specifically to malware. Furthermore, the paper discusses the tool suite, its applications, and how it compares or can work with other related tools.

Peterson, Elena S.↗

jialiu232/MetaFunPrimer_paper_info

Genes belonging to the same functional group may include numerous and variable gene sequences, making characterizing and quantifying difficult. Therefore, high-throughput design tools are needed to simultaneously create primers for improved quantification of target genes. We developed MetaFunPrimer, a bioinformatic pipeline, to design primers for numerous genes of interest. This tool also enables gene target prioritization based on ranking the presence of genes in user-defined references, such as environment-specific metagenomes. Given inputs of protein and nucleotide sequences for gene targets of interest and an accompanying set of reference metagenomes or genomes, MetaFunPrimer generates primers for ranked genes of interest. To demonstrate the usage and benefits of MetaFunPrimer, a total of 78 primer pairs were designed to target observed ammonia monooxygenase subunit A (amoA) genes of ammonia-oxidizing bacteria (AOB) in 1,550 publicly available soil metagenomes. We demonstrate computationally that these amoA-AOB primers can cover 94% of the amoA-AOB genes observed in the 1,550 soil metagenomes compared with a 49% estimated coverage by previously published primers. Finally, we verified the utility of these primer sets in incubation experiments that used long-term nitrogen fertilized or unfertilized soils. High-throughput quantitative PCR (qPCR) results and statistical analyses showed significant differences in relative quantification patterns between the two soils, and subsequent absolute quantifications also confirmed that target genes enumerated by six selected primer pairs were significantly more abundant in the nitrogen-fertilized soils. This new tool gives microbial ecologists a new approach to assess functional gene abundance and related microbial community dynamics quickly and affordably.

Liu, Jia↗