Engineering PapersSearch

SEARCH · Engineering Papers

Results for “bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Measurements of soil protist richness and community composition are influenced by primer pair, annealing temperature, and bioinformatics choices

ABSTRACT Protists are a diverse and understudied group of microbial eukaryotic organisms especially in terrestrial environments. Advances in molecular methods are increasing our understanding of the distribution and functions of these creatures; however, there is a vast array of choices researchers make including barcoding genes, primer pairs, PCR settings, and bioinformatic options that can impact the outcome of protist community surveys. Here, we tested four commonly used primer pairs targeting the V4 and V9 regions of the 18S rRNA gene using different PCR annealing temperatures and processed the sequences with different bioinformatic parameters in 10 diverse soils to evaluate how primer pair, amplification parameters, and bioinformatic choices influence the composition and richness of protist and non-protist taxa using Illumina sequencing. Our results showed that annealing temperature influenced sequencing depth and protist taxon richness for most primer pairs, and that merging forward and reverse sequencing reads for the V4 primer pairs dramatically reduced the number of sequences and taxon richness of protists. The data sets of primers that targeted the same 18S rRNA gene region (e.g., V4 or V9) had similar protist community compositions; however, data sets from primers targeting the V4 18S rRNA gene region detected a greater number of protist taxa compared to those prepared with primers targeting the V9 18S rRNA region. There was limited overlap of protist taxa between data sets targeting the two different gene regions (80/549 taxa). Together, we show that laboratory and bioinformatic choices can substantially affect the results and conclusions about protist diversity and community composition using metabarcoding. IMPORTANCE Ecosystem functioning is driven by the activity and interactions of the microbial community, in both aquatic and terrestrial environments. Protists are a group of highly diverse, mostly unicellular microbes whose identity and roles in terrestrial ecosystem ecology have been largely ignored until recently. This study highlights the importance of choices researchers make, such as primer pair, on the results and conclusions about protist diversity and community composition in soils. In order to better understand the roles protist taxa play in terrestrial ecosystems, biases in methodological and analytical choices should be understood and acknowledged.

Biotechnology & Applied Microbiology

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing

Abstract The microbiome is a complex community of microorganisms, encompassing prokaryotic (bacterial and archaeal), eukaryotic, and viral entities. This microbial ensemble plays a pivotal role in influencing the health and productivity of diverse ecosystems while shaping the web of life. However, many software suites developed to study microbiomes analyze only the prokaryotic community and provide limited to no support for viruses and microeukaryotes. Previously, we introduced the Viral Eukaryotic Bacterial Archaeal (VEBA) open-source software suite to address this critical gap in microbiome research by extending genome-resolved analysis beyond prokaryotes to encompass the understudied realms of eukaryotes and viruses. Here we present VEBA 2.0 with key updates including a comprehensive clustered microeukaryotic protein database, rapid genome/protein-level clustering, bioprospecting, non-coding/organelle gene modeling, genome-resolved taxonomic/pathway profiling, long-read support, and containerization. We demonstrate VEBA’s versatile application through the analysis of diverse case studies including marine water, Siberian permafrost, and white-tailed deer lung tissues with the latter showcasing how to identify integrated viruses. VEBA represents a crucial advancement in microbiome research, offering a powerful and accessible software suite that bridges the gap between genomics and biotechnological solutions.

59 BASIC BIOLOGICAL SCIENCES

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

A Tale from the Trenches: Applying Metamorphic and Differential Testing to Bioinformatics Software

Metamorphic and differential testing have been proposed as best practices for testing software that is difficult to test, such as for programs in scientific domains. An assumption is that these approaches can be easily customized and applied to almost any domain. However, scientific software is often data-driven, and metamorphic relations may require significant domain knowledge to develop. In addition, tools are often written for ad-hoc experimentation by the scientists and often embed many assumptions about the importance and representation of different natural phenomena. In this paper, we present our experience applying both metamorphic and differential testing to a set of four computational biology tools that predict the growth of an organism. While our original goal was to evaluate these techniques to improve our system-level testing, we encountered multiple roadblocks along the way. Although we did find faults (some confirmed by developers), we also uncovered a set of challenges, including the considerable manual effort required for (a) defining domain-specific tests, (b) validating correctness, and (c) distinguishing between issues stemming from poor data and those arising from incorrect software.

Marsh, Alexis L [Iowa State University/Ames Labora

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES

Beneath the surface: Unsolved questions in soil virus ecology

Soil virus ecology is an exciting but still nascent field of research in soil microbiology. While there has been a recent surge in soil virus research studies, many fundamental questions remain unanswered, and a range of technical and bioinformatic challenges need to be overcome. In this perspective article, we present a series of key questions that highlight fruitful research areas for ongoing and future efforts. These include describing the challenges involved in understanding soil viral abundance and activity, spatiotemporal dynamics, life strategy prevalence, virus-mediated biogeochemical impacts, viral protein function, host prediction, and soil RNA virus discovery. In the near term, combining approaches (e.g., cultivation-based, meta-omics, biogeochemical, experimental, and bioinformatic) will be key to assessing the ecological and biogeochemical impacts of soil viruses from the microscopic to the field and global scales. Still, we stress that results must be tempered by current methodological limitations and highlight knowledge gaps that are most pressing to fill via new methods or measurements, such as the prevalence of different viral replication strategies across soils, the fate of microbial necromass carbon after viral lysis, the frequency of virus-host encounters that do not lead to successful infections yet could be bioinformatically mistaken as infections, and the diversity and ecological impacts of RNA viruses in soil.

59 BASIC BIOLOGICAL SCIENCES

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)

Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high-throughput identification and lineage assignment in Pakistan

Bacillus anthracis, the causative agent of anthrax, is a highly virulent zoonotic pathogen primarily affecting domesticated and wild herbivores. Human exposure to B. anthracis is primarily through contact with infected animals or contaminated animal products. In Pakistan, where livestock vaccines are largely unavailable and infected carcasses are often disposed of improperly, the risk to humans, wildlife and livestock is significant. Currently, the diagnosis of anthrax infections and outbreak tracing necessitates the isolation and culturing of B. anthracis, a process that requires BSL-3 facilities. In this study, we show that positive identification, genome reconstruction and lineage assignment can be accomplished using bioinformatic analysis of DNA extracted directly from environmental samples that would otherwise provide the starting material for isolation and culturing. This approach does not require laboratory target enrichment as is necessary for other pathogens, due in part to the extremely high bacterial load in the bloodstream in the deceased animals. Using these methods, we greatly expand the knowledge of endemic B. anthracis in Pakistan. We provide the first reference B. anthracis genomes from Pakistan since the 1970s and identify A.Br.014 Aust94 as a minor circulating sublineage alongside the dominant A.Br.047 Vollum. Future work will focus on the limits of detection and will determine if this bioinformatic method can be expanded more broadly for B. anthracis or other pathogens to replace typical culture-based methods.

A.Br.047 Vollum

Xylanolytic metabolism is regulated by coordination of transcription factors XynR and XylR in extremely thermophilic Caldicellulosiruptorales

ABSTRACT Global transcription factors (TFs) control metabolic processes in bacteria to efficiently utilize available carbon. The orderCaldicellulosiruptoraleshas drawn interest due to the ability of its members to degrade components of lignocellulosic biomass. Regulatory reconstruction ofAnaerocellum (f. Caldicellulosiruptor) besciiidentified two major global transcription factors for xylan utilization, XynR and XylR, and the corresponding putative transcription factor binding sites. Recombinant versions of XynR (LacI family) and XylR (ROK family) were subjected to fluorescence polarization (FP) and biolayer interferometry (BLI) analysis to confirm the predicted binding sites. Four XynR sites and two XylR sites were validated, accounting for 20 of 26 genes regulated by XynR and six of seven genes regulated by XylR. Bioinformatic analysis of the individual genes controlled by the two regulators showed an inter-dependent scheme for xylan conversion; the transport of xylooligosaccharides (XOS) is dependent on XylR, while enzymes responsible for hydrolysis are controlled by both regulators. For xylose catabolism by the xylose isomerase-xylulose kinase pathway, regulation is also split, with XylR controlling xylose isomerase and XynR controlling xylokinase. The XynR/XylR regulator pair withinA. besciiis conserved in all sequenced species ofCaldicellulosiruptorales, suggesting similarities in regulating linear xylan conversion. In other xylanolytic thermophiles, XylR homologs control xylan degradation, compared to just 6 out of 26 genes forA. bescii. These results show that two separate regulatory schemes (dual repression) are coordinated byA. besciito effectively regulate the hemicellulose inventory and xylan catabolism. IMPORTANCE To take full advantage of extreme thermophiles as platform metabolic engineering microorganisms, the tools for genetic manipulation must be further developed, and strategies that exploit a better understanding of metabolic regulation need to be discerned.Anaerocellum bescii, the most studied of the extremely thermophilic fermentative anaerobic bacteria that can utilize microcrystalline cellulose, can degrade microcrystalline cellulose and hemicellulose and has been metabolically engineered to convert the resulting sugars to products such as ethanol and acetone. For xylan, in particular, two major global transcription factors (TFs), XynR and XylR, play a role in sugar metabolism, although their predicted regulatory interdependence from bioinformatics analysis has not been elucidated experimentally. Here, fluorescence polarization (FP) and biolayer interferometry (BLI) were used to explore this issue to support metabolic engineering efforts aimed at improving carbohydrate processing to industrial chemicals.

Biotechnology & Applied Microbiology

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES