Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bioinformatics tool”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Natural Product Gene Clusters in the Filamentous Nostocales Cyanobacterium HT-58-2

Cyanobacteria are known as rich repositories of natural products. One cyanobacterial-microbial consortium (isolate HT-58-2) is known to produce two fundamentally new classes of natural products: the tetrapyrrole pigments tolyporphins A–R, and the diterpenoid compounds tolypodiol, 6-deoxytolypodiol, and 11-hydroxytolypodiol. The genome (7.85 Mbp) of the Nostocales cyanobacterium HT-58-2 was annotated previously for tetrapyrrole biosynthesis genes, which led to the identification of a putative biosynthetic gene cluster (BGC) for tolyporphins. Here, bioinformatics tools have been employed to annotate the genome more broadly in an effort to identify pathways for the biosynthesis of tolypodiols as well as other natural products. A putative BGC (15 genes) for tolypodiols has been identified. Four BGCs have been identified for the biosynthesis of other natural products. Two BGCs related to nitrogen fixation may be relevant, given the association of nitrogen stress with production of tolyporphins. The results point to the rich biosynthetic capacity of the HT-58-2 cyanobacterium beyond the production of tolyporphins and tolypodiols.

60 APPLIED LIFE SCIENCES↗

Application of transport-based metric for continuous interpolation between cryo-EM density maps

Cryogenic electron microscopy (cryo-EM) has become widely used for the past few years in structural biology, to collect single images of macromolecules "frozen in time". As this technique facilitates the identification of multiple conformational states adopted by the same molecule, a direct product of it is a set of 3D volumes, also called EM maps. To gain more insights on the possible mechanisms that govern transitions between different states, and hence the mode of action of a molecule, we recently introduced a bioinformatic tool that interpolates and generates morphing trajectories joining two given EM maps. This tool is based on recent advances made in optimal transport, that allow efficient evaluation of Wasserstein barycenters of 3D shapes. As the overall performance of the method depends on various key parameters, including the sensitivity of the regularization parameter, we performed various numerical experiments to demonstrate how MorphOT can be applied in different contexts and settings. Finally, we discuss current limitations and further potential connections between other optimal transport theories and the conformational heterogeneity problem inherent with cryo-EM data.

3D shapes↗

WHONDRS-GUI: a web application for global survey of surface water metabolites

Background The Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) is a consortium that aims to understand complex hydrologic, biogeochemical, and microbial connections within river corridors experiencing perturbations such as dam operations, floods, and droughts. For one ongoing WHONDRS sampling campaign, surface water metabolite and microbiome samples are collected through a global survey to generate knowledge across diverse river corridors. Metabolomics analysis and a suite of geochemical analyses have been performed for collected samples through the Environmental Molecular Sciences Laboratory (EMSL). The obtained knowledge and data package inform mechanistic and data-driven models to enhance predictions of outcomes of hydrologic perturbations and watershed function, one of the most critical components in model-data integration. To support efforts of the multi-domain integration and make the ever-growing data package more accessible for researchers across the world, a Shiny/R Graphical User Interface (GUI) called WHONDRS-GUI was created. Results The web application can be run on any modern web browser without any programming or operational system requirements, thus providing an open, well-structured, discoverable dataset for WHONDRS. Together with a context-aware dynamic user interface, the WHONDRS-GUI has functionality for searching, compiling, integrating, visualizing and exporting different data types that can easily be used by the community. The web application and data package are available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1484811 , which enables users to simultaneously obtain access to the data and code and to subsequently run the web app locally. The WHONDRS-GUI is also available for online use at Shiny Server ( https://xmlin.shinyapps.io/whondrs/ ).

59 BASIC BIOLOGICAL SCIENCES↗

MultiPhATE2: code for functional annotation and comparison of phage genomes

To address a need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of multiPhATE, a functional annotation code released previously, multiPhATE2 performs gene finding using multiple algorithms, compares the results of the algorithms, performs functional annotation of coding sequences, and incorporates additional search algorithms and databases to extend the search space of the original code. MultiPhATE2 performs gene matching among sets of closely related bacteriophage genomes, and uses multiprocessing to speed computations. MultiPhATE2 can be re-started at multiple points within the workflow to allow the user to examine intermediate results and adjust the subsequent computations accordingly. In addition, multiPhATE2 accommodates custom gene calls and sequence databases, again adding flexibility. MultiPhATE2 was implemented in Python 3.7 and runs as a command-line code under Linux or MAC operating systems. Full documentation is provided as a README file and a Wiki website.

59 BASIC BIOLOGICAL SCIENCES↗

Biological Parts Search Portal (BioParts) v1.0.0

BioParts is a web based search portal for biological parts available in the public domain. It combines the ease and convenience of modern web search engines with the capabilities of bioinformatics search tools such as BLAST. This portal, available at bioparts.org, allows anyone to search for publicly accessible biological part information (e.g., NCBI, iGEM, SynBioHub, Addgene), including parts publicly accessible through ICE Registries. Additionally, the portal offers a REST API that enables third-party applications and tools to access the portal's functionality programmatically. While there are several standalone biological part repositories, there doesn't exist an application that indexes these publicly available parts and enables features such as keyword and BLAST searches along with automatic sequence annotation.

Plahar, Hector↗

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat↗

How Cooperative Engagement Programs Strengthen Sequencing Capabilities for Biosurveillance and Outbreak Response

The threat of emerging and re-emerging infectious diseases continues to be a challenge to public and global health security. Cooperative biological engagement programs act to build partnerships and collaborations between scientists and health professionals to strengthen capabilities in biosurveillance. Biosurveillance is the systematic process of detecting, reporting, and responding to especially dangerous pathogens and pathogens of pandemic potential before they become outbreaks, epidemics, and pandemics. One important tool in biosurveillance is next generation sequencing. Expensive sequencing machines, reagents, and supplies make it difficult for countries to adopt this technology. Cooperative engagement programs help by providing funding for technical assistance to strengthen sequencing capabilities. Through workshops and training, countries are able to learn sequencing and bioinformatics, and implement these tools in their biosurveillance programs. Cooperative programs have an important role in building and sustaining collaborations among institutions and countries. One of the most important pieces in fostering these collaborations is trust. Trust provides the confidence that a successful collaboration will benefit all parties involved. With sequencing, this enables the sharing of pathogen samples and sequences. Obtaining global sequencing data helps to identify unknown etiological agents, track pathogen evolution and infer transmission networks throughout the duration of a pandemic. Having sequencing technology in place for biosurveillance generates the capacity to provide real-time data to understand and respond to pandemics. We highlight the need for these programs to continue to strengthen sequencing in biosurveillance. By working together to strengthen sequencing capabilities, trust can be formed, benefitting global health in the face of biological threats.

60 APPLIED LIFE SCIENCES↗

Appendix Q: Recommendations for Developing Molecular Assays for Microbial Pathogen Detection Using Modern In Silico Approaches

We describe the use of in silico approaches to improve the process of molecular assay development and reduce time and cost by utilizing available databases of whole genome pathogen sequences combined with modern bioinformatics and physical modeling tools. Well-characterized assays are needed for accurately detecting pathogens in environmental and patient samples and also for evaluation of the efficacy of a medical countermeasure that may be administered to patients. The polymerase chain reaction (PCR) remains the gold standard for pathogen detection due to the simplicity of its instrumentation, low cost of reagents, and outstanding limit of detection (LOD), sensitivity, and specificity. However, creation of such PCR assays often involves iterations of design, preliminary testing, and thorough validation with clinical isolates and testing in relevant matrices, which can be time consuming, costly, and result in suboptimal assays. Since formal validation (e.g., for Emergency Use Authorization [EUA] or Food and Drug Administration [FDA] licensure) of an infectious disease assay can be very expensive and can require extensive time of development, having a well-designed assay up front is a critical first step. Yet, many assays described in the literature utilized limited design capabilities and many initially promising assays fail the validation process, resulting in increased costs and timelines for successful product development. While the computational approaches outlined in this document by no means obviate the need for wet lab testing, they can reduce the amount of effort wasted on empirical optimization and iterative redesigns and also guide validation studies. The proposed computational approaches also result in higher performing assays with better sensitivity, specificity, and lower LOD and reduce the possibility of assay failure due to signature erosion. To provide clarity, an extensive glossary of defined terms is provided.

59 BASIC BIOLOGICAL SCIENCES↗

AlphaBeta: computational inference of epimutation rates and spectra from high-throughput DNA methylation data in plants

Stochastic changes in DNA methylation (i.e., spontaneous epimutations) contribute to methylome diversity in plants. Here, we describe AlphaBeta, a computational method for estimating the precise rate of such stochastic events using pedigree-based DNA methylation data as input. We demonstrate how AlphaBeta can be employed to study transgenerationally heritable epimutations in clonal or sexually derived mutation accumulation lines, as well as somatic epimutations in long-lived perennials. Application of our method to published and new data reveals that spontaneous epimutations accumulate neutrally at the genome-wide scale, originate mainly during somatic development and that they can be used as a molecular clock for age-dating trees.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Global overview and major challenges of host prediction methods for uncultivated phages

Bacterial communities play critical roles across all of Earth’s biomes, affecting human health and global ecosystem functioning. They do so under strong constraints exerted by viruses, i.e., bacteriophages or “phages”. Phages can reshape bacterial communities’ structure, influence long-term evolution of bacterial populations, and alter host cell metabolism during infection. Metagenomics approaches, i.e., shotgun sequencing of environmental DNA or RNA, recently enabled large-scale exploration of phage genomic diversity, yielding several millions of phage genomes now to be further analyzed and characterized. One major challenge however is the lack of direct host information for these phages. Several methods and tools have been proposed to bioinformatically predict the potential host(s) of uncultivated phages based only on genome sequence information. Here we review these different approaches and highlight their distinct strengths and limitations. We also outline complementary experimental assays which are being proposed to validate and refine these bioinformatic predictions.

59 BASIC BIOLOGICAL SCIENCES↗

Flexible and Adaptive Malware Identification Using Techniques from Biology

The holy grail in cyber analytics is to find new ways to understand the information we already have access to. One way to do that is to characterize the data into reasonable sizes and then leverage any known information to generate new insights. Biologists have been using a similar process for decades. This paper introduces the MLSTONES tool set that was developed by leveraging biology and bioinformatics, high performance computing, and statistical algorithms applied to cyber data and specifically to malware. Furthermore, the paper discusses the tool suite, its applications, and how it compares or can work with other related tools.

Peterson, Elena S.↗

jialiu232/MetaFunPrimer_paper_info

Genes belonging to the same functional group may include numerous and variable gene sequences, making characterizing and quantifying difficult. Therefore, high-throughput design tools are needed to simultaneously create primers for improved quantification of target genes. We developed MetaFunPrimer, a bioinformatic pipeline, to design primers for numerous genes of interest. This tool also enables gene target prioritization based on ranking the presence of genes in user-defined references, such as environment-specific metagenomes. Given inputs of protein and nucleotide sequences for gene targets of interest and an accompanying set of reference metagenomes or genomes, MetaFunPrimer generates primers for ranked genes of interest. To demonstrate the usage and benefits of MetaFunPrimer, a total of 78 primer pairs were designed to target observed ammonia monooxygenase subunit A (amoA) genes of ammonia-oxidizing bacteria (AOB) in 1,550 publicly available soil metagenomes. We demonstrate computationally that these amoA-AOB primers can cover 94% of the amoA-AOB genes observed in the 1,550 soil metagenomes compared with a 49% estimated coverage by previously published primers. Finally, we verified the utility of these primer sets in incubation experiments that used long-term nitrogen fertilized or unfertilized soils. High-throughput quantitative PCR (qPCR) results and statistical analyses showed significant differences in relative quantification patterns between the two soils, and subsequent absolute quantifications also confirmed that target genes enumerated by six selected primer pairs were significantly more abundant in the nitrogen-fertilized soils. This new tool gives microbial ecologists a new approach to assess functional gene abundance and related microbial community dynamics quickly and affordably.

Liu, Jia↗

GNET2: an R package for constructing gene regulatory networks from transcriptomic data

Abstract Motivation The Gene Network Estimation Tool (GNET) is designed to build gene regulatory networks (GRNs) from transcriptomic gene expression data with a probabilistic graphical model. The data preprocessing, model construction and visualization modules of the original GNET software were developed on different programming platforms, which were inconvenient for users to deploy and use. Results Here, we present GNET2, an improved implementation of GNET as an integrated R package. GNET2 provides more flexibility for parameter initialization and regulatory module construction based on the core iterative modeling process of the original algorithm. The data exchange interface of GNET2 is handled within an R session automatically. Given the growing demand for regulatory network reconstruction from transcriptomic data, GNET2 offers a convenient option for GRN inference on large datasets. Availability and implementation The source code of GNET2 is available at https://github.com/jianlin-cheng/GNET2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

RCSB Protein Data Bank: Celebrating 50 years of the PDB with new tools for understanding and visualizing biological macromolecules in 3D

We report the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB), funded by the US National Science Foundation, National Institutes of Health, and Department of Energy, has served structural biologists and Protein Data Bank (PDB) data consumers worldwide since 1999. RCSB PDB, a founding member of the Worldwide Protein Data Bank (wwPDB) partnership, is the US data center for the global PDB archive housing biomolecular structure data. RCSB PDB is also responsible for the security of PDB data, as the wwPDB-designated Archive Keeper. Annually, RCSB PDB serves tens of thousands of three-dimensional (3D) macromolecular structure data depositors (using macromolecular crystallography, nuclear magnetic resonance spectroscopy, electron microscopy, and micro-electron diffraction) from all inhabited continents. RCSB PDB makes PDB data available from its research-focused RCSB.org web portal at no charge and without usage restrictions to millions of PDB data consumers working in every nation and territory worldwide. In addition, RCSB PDB operates an outreach and education PDB101.RCSB.org web portal that was used by more than 800,000 educators, students, and members of the public during calendar year 2020. This invited Tools Issue contribution describes (i) how the archive is growing and evolving as new experimental methods generate ever larger and more complex biomolecular structures; (ii) the importance of data standards and data remediation in effective management of the archive and facile integration with more than 50 external data resources; and (iii) new tools and features for 3D structure analysis and visualization made available during the past year via the RCSB.org web portal.

59 BASIC BIOLOGICAL SCIENCES↗