Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “lipid annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Preprocessing Tool for Enhanced Ion Mobility–Mass Spectrometry-Based Omics Workflows

The ability to improve the data quality of ion mobility–mass spectrometry (IM-MS) measurements is of great importance for enabling modular and efficient computational workflows and gaining better qualitative and quantitative insights from complex biological and environmental samples. We developed the PNNL PreProcessor, a standalone and user-friendly software housing various algorithmic implementations to generate new MS-files with enhanced signal quality and in the same instrument format. Different experimental approaches are supported for IM-MS based on Drift-Tube (DT) and Structures for Lossless Ion Manipulations (SLIM), including liquid chromatography (LC) and infusion analyses. The algorithms extend the dynamic range of the detection system, while reducing file sizes for faster and memory-efficient downstream processing. Specifically, multidimensional smoothing improves peak shapes of poorly defined low-abundance signals, and saturation repair reconstructs the intensity profile of high-abundance peaks from various analyte types. Further, other functionalities are data compression and interpolation, IM demultiplexing, noise filtering by low intensity threshold and spike removal, and exporting of acquisition metadata. Several advantages of the tool are illustrated, including an increase of 19.4% in lipid annotations and a two-times faster processing of LC-DT IM-MS data-independent acquisition spectra from a complex lipid extract of a standard human plasma sample. The software is freely available at https://omics.pnl.gov/software/pnnl-preprocessor.

59 BASIC BIOLOGICAL SCIENCES↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

Near-complete genome sequence of Lipomyces tetrasporous NRRL Y-64009, an oleaginous yeast capable of growing on lignocellulosic hydrolysates

ABSTRACT Lipomyces tetrasporous is an oleaginous yeast that can utilize a variety of plant-based sugars. It accumulates lipids during growth on lignocellulosic biomass hydrolysates. We present the annotated genome sequence of L. tetrasporous NRRL Y-64009 to aid in its development as a platform organism for producing lipids and lipid-based bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Are Phosphatidic Acids Ubiquitous in Mammalian Tissues or Overemphasized in Mass Spectrometry Imaging Applications?

Abstract Mass spectrometry imaging (MSI) is an invaluable tool for the spatial visualization of molecules in vivo. However, the question of whether observed annotations are endogenous or artificial (i. e., from in‐source fragmentation) is critical and has been largely unexplored in multimodal MSI. In matrix‐assisted laser desorption/ionization (MALDI)‐MSI datasets from researchers worldwide, PAs were found to represent up to 18 % of annotations in rat brain. Rat brain was additionally imaged here using nanospray desorption electrospray ionization (nano‐DESI), a softer ionization strategy. No PAs observed with MALDI were present in the nano‐DESI dataset. Further investigation strongly indicated lipid fragmentation to PAs for MALDI‐MSI, but not with nano‐DESI‐MSI. We finally extend this observation to the MALDI‐MSI analyses of human tissues, showing that PA annotations comprised up to 16 % of annotations. Therefore, this study shows that MSI annotations should be carefully interrogated, as in‐source fragmentation or modification of lipids may contribute substantially to false annotations and incorrect biological interpretations.

Vandergrift, Gregory W.↗

Root exudate lipids: Uncovering chemodiversity and carbon stability potential

Root-derived carbon has been shown to contribute more to soil carbon stocks than aboveground litter. Yet the molecular chemodiversity of root exudates remains poorly understood due to limited characterization and annotation. In this study, we characterized the molecular chemodiversity and production of metabolites and lipids in root exudates from field grown mature tall wheatgrass (Thinopyrum ponticum). We discovered a diversity of lipids, including substantial levels of triacylglycerols (∼19 μg/g fresh root per min), fatty acyls, sphingolipids, sterol lipids, and glycerophospholipids, some of which have not been previously documented in root exudates. By integrating tandem mass spectral library searching and deep learning-based chemical class assignment, our metabo-lipidomics approach significantly expanded the known molecular diversity of root exudates. Rates of lipid derived carbon production were approximately double that of polar metabolites (lipids: 81.52 ± 13.81 vs polar metabolites: 38.41 ± 5.93 μg C g −1 fresh root mass min −1 ) with an order of magnitude higher carbon to nitrogen ratios (lipids: 459 ± 90 vs polar metabolites: 14.40 ± 0.58). Exudate lipids displayed highly negative nominal oxidation state of carbon (−1.182 to −1.909), indicating that these compounds may be less favorable for microbial decomposition. Together our results suggest the potential of root exudate lipids to contribute to stable carbon pools in soil, supporting long-term carbon storage. This work advances understanding of plant-derived lipid inputs to soil and underscores the need for future studies on the functional roles of lipids in shaping root-microbe-soil interactions, microbial activity, soil structure, and nutrient availability – contributing to soil health.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genome-scale model development and genomic sequencing of the oleaginous clade Lipomyces

The Lipomyces clade contains oleaginous yeast species with advantageous metabolic features for biochemical and biofuel production. Limited knowledge about the metabolic networks of the species and limited tools for genetic engineering have led to a relatively small amount of research on the microbes. Here, a genome-scale metabolic model (GSM) of Lipomyces starkeyi NRRL Y-11557 was built using orthologous protein mappings to model yeast species. Phenotypic growth assays were used to validate the GSM (66% accuracy) and indicated that NRRL Y-11557 utilized diverse carbohydrates but had more limited catabolism of organic acids. The final GSM contained 2,193 reactions, 1,909 metabolites, and 996 genes and was thus named iLst996. The model contained 96 of the annotated carbohydrate-active enzymes. iLst996 predicted a flux distribution in line with oleaginous yeast measurements and was utilized to predict theoretical lipid yields. Twenty-five other yeasts in the Lipomyces clade were then genome sequenced and annotated. Sixteen of the Lipomyces species had orthologs for more than 97% of the iLst996 genes, demonstrating the usefulness of iLst996 as a broad GSM for Lipomyces metabolism. Pathways that diverged from iLst996 mainly revolved around alternate carbon metabolism, with ortholog groups excluding NRRL Y-11557 annotated to be involved in transport, glycerolipid, and starch metabolism, among others. Overall, this study provides a useful modeling tool and data for analyzing and understanding Lipomyces species metabolism and will assist further engineering efforts in Lipomyces .

59 BASIC BIOLOGICAL SCIENCES↗

Experimental Assessment of Mammalian Lipidome Complexity Using Multimodal 21 T FTICR Mass Spectrometry Imaging

Herein, we assess the complementarity and complexity of data that can be detected within mammalian lipidome mass spectrometry imaging (MSI) via matrix-assisted laser desorption ionization (MALDI) and nanospray desorption electrospray ionization (nano-DESI). We do so by employing 21 T Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) with absorption mode FT processing in both cases, allowing unmatched mass resolving power per unit time (≥613k at m/z 760, 1.536 s transients). While our results demonstrated that molecular coverage and dynamic range capabilities were greater in MALDI analysis, nano-DESI provided superior mass error, and all annotations for both modes had sub-ppm error. Taken together, these experiments highlight the coverage of 1676 lipids and serve as a functional guide for expected lipidome complexity within nano-DESI-MSI and MALDI-MSI. To further assess the lipidome complexity, mass splits (i.e., the difference in mass between neighboring peaks) within single pixels were collated across all pixels from each respective MSI experiment. The spatial localization of these mass splits was powerful in informing whether the observed mass splits were biological or artificial (e.g., matrix related). Mass splits down to 2.4 mDa were observed (i.e., sodium adduct ambiguity) in each experiment, and both modalities highlighted comparable degrees of lipidome complexity. Further, we highlight the persistence of certain mass splits (e.g., 8.9 mDa; double bond ambiguity) independent of ionization biases. Here, we also evaluate the need for ultrahigh mass resolving power for mass splits ≤4.6 mDa (potassium adduct ambiguity) at m/z > 1000, which may only be resolved by advanced FTICR-MS instrumentation.

21T-FTICR-MS↗

IsoMatchMS : Open-Source Software for Automated Annotation and Visualization of High Resolution MALDI-MS Spectra

Due to its speed, accuracy, and adaptability to various sample types, matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) has become a popular method to identify molecular isotope profiles from biological samples. Often MALDI-MS data do not include tandem MS fragmentation data, and thus the identification of compounds in samples requires external databases so that the accurate mass of detected signals can be matched to known molecular compounds. Most relevant MALDI-MS software tools developed to confirm compound identifications are focused on small molecules (e.g., metabolites, lipids) and cannot be easily adapted to protein data due to their more complex isotopic distributions. Here, we present an R package called IsoMatchMS for the automated annotation of MALDI-MS data for multiple datatypes: intact proteins, peptides, and glycans. This tool accepts already derived molecular formulas or, for proteomics applications, can derive molecular formulas from a list of input peptides or proteins including proteins with post-translational modifications. In conclusion, visualization of all matched isotopic profiles is provided in a highly accessible HTML format called a trelliscope display, which allows users to filter and sort by several parameters such as match scores and the number of peaks matched. IsoMatchMS simplifies the annotation and visualization of MALDI-MS data for downstream analyses.

47 OTHER INSTRUMENTATION↗

Genome sequence, phylogenetic analysis, and structure-based annotation reveal metabolic potential of Chlorella sp. SLA-04

Algae are a broad class of photosynthetic eukaryotes that are phylogenetically and physiologically diverse. Most of the phylogenetic diversity has been inferred from 18S rDNA sequencing since there are only a few complete genomes available in public databases. Here we use ultra-long-read Nanopore sequencing to determine a gapless, telomere-to-telomere complete genome sequence of Chlorella sp. SLA-04, previously described as Chlorella sorokiniana SLA-04. Chlorella sp. SLA-04 is a green alga that grows to high cell density in a wide variety of environments - high and neutral pH, high and low alkalinity, and high and low salinity. SLA-04's ability to grow in high pH and high alkalinity media without external CO 2 supply is favorable for large-scale algal biomass production. Phylogenetic analysis performed using ribosomal DNA and conserved protein sequences consistently reveal that Chlorella sp. SLA-04 forms a distinct lineage from other strains of Chlorella sorokiniana. We complement traditional genome annotation methods with high throughput structural predictions and demonstrate that this approach expands functional prediction of the SLA-04 proteome. Genomic analysis of the SLA-04 genome identifies the genes capable of utilizing TCA cycle intermediates to replenish cytosolic acetyl-CoA pools for lipid production. We also identify a complete metabolic pathway for sphingolipid anabolism that may allow SLA-04 to readily adapt to changing environmental conditions and facilitate robust cultivation in mass production systems. Altogether, this work clarifies the phylogeny of Chlorella sp. SLA-04 within Trebouxiophyceae and demonstrates how structural predictions can be used to improve annotation beyond sequencebased methods.

59 BASIC BIOLOGICAL SCIENCES↗

Atomistic simulations for investigation of substrate and salt effects on lipid in-source fragmentation in secondary ion mass spectrometry: A follow-up study

In-source fragmentation (ISF) poses a significant challenge in secondary ion mass spectrometry (SIMS). These fragment ions increase the spectral complexity and can lead to incorrect annotation of fragments as intact species. The presence of salt that is ubiquitous in biological samples can influence the fragmentation and ionization of analytes in a significant manner, but their influences on SIMS have not been well characterized. To elucidate the effect of substrates and salt on ISF in SIMS, we have employed experimental SIMS in combination with atomistic simulations of a sphingolipid on a gold surface with various NaCl concentrations as a model system. Our results revealed that a combination of bond dissociation energy and binding energy between N-palmitoyl-sphingomyelin and a gold surface is a good predictor of fragment ion intensities in the absence of salt. However, ion-fragment interactions play a significant role in determining fragment yields in the presence of salt. Additionally, the charge distribution on fragment species may be a major contributor to the varying effects of salt on fragmentation. This study demonstrates that atomistic modeling can help predict ionization potential when salts are present, providing insights for more accurate interpretations of complex biological spectra.

74 ATOMIC AND MOLECULAR PHYSICS↗

Actinomycetota isolated from the sponge Hymeniacidon perlevis as a source of novel compounds with pharmacological applications: diversity, bioactivity screening, and metabolomic analysis

Abstract Aims To combat health conditions, such as multi-resistant bacterial infections, cancer, and metabolic diseases, new drugs need to be urgently found and, in this respect, marine Actinomycetota have a high potential to produce secondary metabolites with pharmacological importance. We aimed to study the cultivable Actinomycetota community associated with a marine sponge from the Portuguese coast, Hymeniacidon perlevis, and investigate the potential of the retrieved isolates to produce compounds with antimicrobial, anticancer and anti-obesity properties. Methods and results The analysis of the 16S rRNA gene revealed 79 Actinomycetota isolates affiliated with 12 genera—Brachybacterium, Dietzia, Glutamicibacter, Gordonia, Micrococcus, Micromonospora, Nocardia, Nocardiopsis, Paenoartrhobacter, Rhodococcus, Streptomyces, and Tsukamurella, most of which affiliated with the genus Streptomyces. The screening of antimicrobial activity revealed 13 strains, all belonging to the Streptomyces genus, capable of inhibiting the growth of Candida albicans, Bacillus subtilis, or Staphylococcus aureus. Forty-three extracts exhibited cytotoxic activity against at least one tested cell line (HepG2, HCT-116, and hCMEC-D3). Three extracts that were active against the two cancer cell lines tested, did not reduce the viability of the non-cancer endothelial cell line, hCMEC-D3. One Gordonia strain exhibited anti-obesity activity, revealed by its ability to reduce the neutral lipids in zebrafish larvae. Mass spectrometry-based dereplication analysis of active extracts identified several compounds associated with known Actinomycetota natural products. Nonetheless, five clusters contained metabolites that did not match any annotated natural products, suggesting they may represent new bioactive molecules. Conclusions This work contributed to increase the knowledge on the diversity and bioactive potential of Actinomycetota associated with H. perlevis.

Fonseca, Ana C.↗

The Transcriptional Response of Soil Bacteria to Long-Term Warming and Short-Term Seasonal Fluctuations in a Terrestrial Forest

Terrestrial ecosystems are an important carbon store, and this carbon is vulnerable to microbial degradation with climate warming. After 30 years of experimental warming, carbon stocks in a temperate mixed deciduous forest were observed to be reduced by 30% in the heated plots relative to the controls. In addition, soil respiration was seasonal, as was the warming treatment effect. We therefore hypothesized that long-term warming will have higher expressions of genes related to carbohydrate and lipid metabolism due to increased utilization of recalcitrant carbon pools compared to controls. Because of the seasonal effect of soil respiration and the warming treatment, we further hypothesized that these patterns will be seasonal. We used RNA sequencing to show how the microbial community responds to long-term warming (~30 years) in Harvard Forest, MA. Total RNA was extracted from mineral and organic soil types from two treatment plots (+5°C heated and ambient control), at two time points (June and October) and sequenced using Illumina NextSeq technology. Treatment had a larger effect size on KEGG annotated transcripts than on CAZymes, while soil types more strongly affected CAZymes than KEGG annotated transcripts, though effect sizes overall were small. Although, warming showed a small effect on overall CAZymes expression, several carbohydrate-associated enzymes showed increased expression in heated soils (~68% of all differentially expressed transcripts). Further, exploratory analysis using an unconstrained method showed increased abundances of enzymes related to polysaccharide and lipid metabolism and decomposition in heated soils. Compared to long-term warming, we detected a relatively small effect of seasonal variation on community gene expression. Together, these results indicate that the higher carbohydrate degrading potential of bacteria in heated plots can possibly accelerate a self-reinforcing carbon cycle-temperature feedback in a warming climate.

54 ENVIRONMENTAL SCIENCES↗

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics ↗

Natural variation and improved genome annotation of the emerging biofuel crop field pennycress ( Thlaspi arvense )

The Brassicaceae family comprises more than 3,700 species with a diversity of phenotypic characteristics, including seed oil content and composition. Recently, the global interest in Thlaspi arvense L. (pennycress) has grown as the seed oil composition makes it a suitable source for biodiesel and aviation fuel production. However, many wild traits of this species need to be domesticated to make pennycress ideal for cultivation. Molecular breeding and engineering efforts require the availability of an accurate genome sequence of the species. Here, we describe pennycress genome annotation improvements, using a combination of long- and short-read transcriptome data obtained from RNA derived from embryos of 22 accessions, in addition to public genome and gene expression information. Our analysis identified 27,213 protein-coding genes, as well as on average 6,188 biallelic SNPs. In addition, we used the identified SNPs to evaluate the population structure of our accessions. The data from this analysis support that the accession Ames 32872, originally from Armenia, is highly divergent from the other accessions, while the accessions originating from Canada and the United States cluster together. When we evaluated the likely signatures of natural selection from alternative SNPs, we found 7 candidate genes under likely recent positive selection. These genes are enriched with functions related to amino acid metabolism and lipid biosynthesis and highlight possible future targets for crop improvement efforts in pennycress.

59 BASIC BIOLOGICAL SCIENCES↗

Interpreting the Lipidome: Bioinformatic Approaches to Embrace the Complexity

Background Improvements in mass spectrometry (MS) technologies coupled with bioinformatics developments have allowed considerable advancement in the measurement and interpretation of lipidomics data in recent years. Since research areas employing lipidomics are rapidly increasing, there is a great need for bioinformatic tools that capture and utilize the complexity of the data. Currently, the diversity and complexity within the lipidome is often concealed by summing over or averaging individual lipids up to (sub)class-based descriptors, losing valuable information about biological function and interactions with other distinct lipids molecules, proteins and/or metabolites. Aim of review To address this gap in knowledge, novel bioinformatics methods are needed to improve identification, quantification, integration and interpretation of lipidomics data. The purpose of this mini-review is to summarize exemplary methods to explore the complexity of the lipidome. Key scientific concepts of review Here we describe six approaches that capture three core focus areas for lipidomics: (1) lipidome annotation including a resolvable database identifier, (2) interpretation via pathway- and enrichment-based methods, and (3) understanding complex interactions to emphasize specific steps in the analytical process and highlight challenges in analyses associated with the complexity of lipidome data.

Kyle, Jennifer E.↗

Metabolomics Analysis of Bacterial Pathogen Burkholderia thailandensis and Mammalian Host Cells in Co-culture

The Tier 1 HHS/USDA Select Agent Burkholderia pseudomallei is a bacterial pathogen that is highly virulent when introduced into the respiratory tract and intrinsically resistant to many antibiotics. Transcriptomic- and proteomic-based methodologies have been used to investigate mechanisms of virulence employed by B. pseudomallei and Burkholderia thailandensis, a convenient surrogate; however, analysis of the pathogen and host metabolomes during infection is lacking. Changes in the metabolites produced can be a result of altered gene expression and/or post-transcriptional processes. Thus, metabolomics complements transcriptomics and proteomics by providing a chemical readout of a biological phenotype, which serves as a snapshot of an organism’s physiological state. However, the poor signal from bacterial metabolites in the context of infection poses a challenge in their detection and robust annotation. In this work, we coupled mammalian cell culture-based metabolomics with feature-based molecular networking of mono- and co-cultures to annotate the pathogen’s secondary metabolome during infection of mammalian cells. These methods enabled us to identify several key secondary metabolites produced by B. thailandensis during infection of airway epithelial and macrophage cell lines. Additionally, the use of in silico approaches provided insights into shifts in host biochemical pathways relevant to defense against infection. Using chemical class enrichment analysis, for example, we identified changes in a number of host-derived compounds including immune lipids such as prostaglandins, which were detected exclusively upon pathogen challenge. Taken together, our findings indicate that co-culture of B. thailandensis with mammalian cells alters the metabolome of both pathogen and host and provides a new dimension of information for in-depth analysis of the host–pathogen interactions underlying Burkholderia infection.

60 APPLIED LIFE SCIENCES↗

A multi-omic characterization of the physiological responses to salt stress in Scenedesmus obliquus UTEX393

Scenedesmus obliquus UTEX393 is a promising microalgal candidate for sustainable biomanufacturing but its limited halotolerance hinders large-scale cultivation in saline environments. To investigate the molecular basis of salt stress responses, we conducted a comprehensive multi-omic analysis integrating genomics, transcriptomics, proteomics, lipidomics, metabolomics, and DNA affinity purification sequencing (DAP-seq). An improved nuclear genome assembly and annotation yielded 19,017 gene models and a 97% BUSCO completeness score, enabling construction of a genome-scale metabolic model. Comparing 15 ppt salinity stress to 5 ppt control, growth and productivity were significantly reduced, accompanied by widespread transcriptomic and proteomic changes. Transcriptomic analysis revealed downregulation of photosynthetic machinery and energy conservation genes, and upregulation of stress-responsive elements such as expansins, flavodoxins, and osmoprotectants. Lipidomic profiling showed accumulation of triacylglycerols (TAGs) and degradation of galactosyl lipids, consistent with a shift toward lipid biosynthesis to mitigate redox imbalance. Depletion of key polar metabolites and branched-chain amino acids suggested a rerouting of central carbon metabolism under stress. DAP-seq identified key transcription factors, including LHY1 and SPL12, that target central metabolic enzymes involved in redox balancing, such as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and malate dehydrogenase (MDH). These findings establish a regulatory-metabolic framework linking redox stress to lipid accumulation and reveal potential engineering targets to enhance salt tolerance. Overall, the multi-omic analysis supports the “overflow” hypothesis, where impaired photosynthesis results in excess reducing equivalents being diverted into TAG synthesis and highlights transcriptional regulators as candidates for improving algal robustness in brackish environments.

09 BIOMASS FUELS↗

BioTransformer 3.0 – A Web Server for Accurately Predicting Metabolic Transformation Products

BioTransformer 3.0 is a freely available web server that supports accurate, rapid and comprehensive in silico metabolism prediction. It combines machine learning approaches with a rule-based system to predict small-molecule metabolism in human tissues, the human gut as well as the external environment (soil and water microbiota). Simply stated, BioTransformer takes a molecular structure as input (SMILES or SDF) and outputs an interactively viewable/sortable table of the predicted metabolites or transformation products (SMILES, PNG images) along with the enzymes that are predicted to be responsible for those reactions and richly annotated downloadable files (CSV and JSON). The entire process typically takes a few seconds. Previous versions of BioTransformer focused exclusively on predicting the metabolism of xenobiotics (such as plant natural products, drugs, cosmetics and other synthetic compounds) using a limited number of pre-defined steps and somewhat limited rule-based methods. BioTransformer 3.0, uses much more sophisticated methods and incorporates new databases, new constraints and new prediction modules to not only more accurately predict the metabolic transformation products of exogenous xenobiotics but also the transformation products of endogenous metabolites, such as amino acids, peptides, carbohydrates, organic acids, and lipids. BioTransformer 3.0 can also support customized sequential combinations of these transformations along with multiple iterations to simulate multi-step human and/or environmental biotransformation events. Performance tests indicate that BioTransformer 3.0 is 40-50% more accurate, much less prone to combinatorial “explosions” and far more comprehensive in terms of metabolite coverage/capabilities than previous versions of BioTransformer.

59 BASIC BIOLOGICAL SCIENCES↗