Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high-throughput automated workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

express: Extensible, high-level workflows for swifter ab initio materials modeling

In this work, we introduce an open-source Julia project, express, an extensible, lightweight, high-throughput, high-level workflow framework that aims to automate ab initio calculations for the materials science community. express is shipped with well-tested workflow templates, including structure optimization, equation of state (EOS) fitting, phonon spectrum (lattice dynamics) calculation, and thermodynamic property calculation in the framework of the quasi-harmonic approximation (QHA). It is designed to be highly modularized so that its components can be reused across various occasions, and customized workflows can be built on top of that. Users can also track the status of workflows in real-time, and rerun failed jobs thanks to the data lineage feature express provides. Finally, two working examples, i.e., all workflows applied to lime and akimotoite, are also presented in the code and this paper.

36 MATERIALS SCIENCE↗

pathSQE : an automated workflow for single-crystal inelastic neutron scattering data processing and analysis

Inelastic neutron scattering (INS) experiments utilizing modern time-of-flight spectrometers enable the comprehensive mapping of the energy (E)- and momentum (Q)-resolved dynamical structure factor of single crystals, probing both the lattice and magnetic excitations. Yet, the large size and complexity of four-dimensional INS data are challenging current analysis workflows, often resulting in an underutilization of the measured information. To help address this issue, this paper introduces new software interfaced with the Mantid framework, pathSQE, designed to streamline the processing, analysis and interpretation of 4D single-crystal INS data. By automating key tasks such as 1D/2D slicing, symmetrization, Brillouin zone folding, data visualization, prioritization and filtering, and comparisons with simulations, pathSQE facilitates and accelerates INS data analysis workflows. Here, this paper outlines the features and implementation and provides several illustrations of the use of pathSQE on data collected on single crystals using direct-geometry time-of-flight spectrometers at the Spallation Neutron Source, including Ge, FeSi, MnO and SnS single-crystal measurements on the ARCS, HYSPEC and CNCS neutron spectrometers. Beyond streamlining post-experiment data processing, pathSQE establishes an automated and modular processing pipeline that could support future real-time experiment steering.

36 MATERIALS SCIENCE↗

Workflow for High-throughput Screening of Enzyme Mutant Libraries Using Matrix-assisted Laser Desorption/Ionization Mass Spectrometry Analysis of Escherichia coli Colonies

High-throughput molecular screening of microbial colonies and DNA libraries are critical procedures that enable applications such as directed evolution, functional genomics, microbial identification, and creation of engineered microbial strains to produce high-value molecules. A promising chemical screening approach is the measurement of products directly from microbial colonies via optically guided matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). Measuring the compounds from microbial colonies bypasses liquid culture with a screen that takes approximately 5 s per sample. We describe a protocol combining a dedicated informatics pipeline and sample preparation method that can prepare up to 3,000 colonies in under 3 h. The screening protocol starts from colonies grown on Petri dishes and then transferred onto MALDI plates via imprinting. The target plate with the colonies is imaged by a flatbed scanner and the colonies are located via custom software. The target plate is coated with MALDI matrix, MALDI-MS analyzes the colony locations, and data analysis enables the determination of colonies with the desired biochemical properties. This workflow screens thousands of colonies per day without requiring additional automation. The wide chemical coverage and the high sensitivity of MALDI-MS enable diverse screening projects such as modifying enzymes and functional genomics surveys of gene activation/inhibition libraries.

Choe, Kisurb↗

Comparison of automated chemical-guided segmentation and human annotation of soil organic matter in X-ray microcomputed tomography imaging in contrasted soil types

Soil organic matter (OM) formation and persistence is strongly influenced by the spatial distribution of organic substrates and microscale soil heterogeneity by dictating OM accessibility to microorganisms. However, traditional size and/or density fractionation techniques disrupt aggregate architecture, eliminating spatial information needed to fully understand intra-aggregate OM distribution. To quantify three-dimensional OM spatial distribution and automate segmentation in X-ray microcomputed tomography (µCT) imaging without human annotation bias, we developed an iodine gas vapor (I2) based staining workflow that eliminates labor-intensive manual annotation while maintaining segmentation accuracy, using aggregates from four taxonomically diverse soils (Xerofluvent, Haploxeroll Sphagnofibrist, Palehumult) with an 8-fold range of soil organic carbon. Human annotation of 10 µCT slices by the experienced and inexperienced annotators resulted in variations up to 3% in the Dice similarity coefficient (DSC), reflecting a degree of inherent subjectivity of manual labeling. Such inconsistencies are expected to compound as the number of manually annotated slices increases. Dual-energy µCT imaging at 33.1 keV (below the iodine (I) K-edge) and 33.2 keV (above the I K-edge) was used to resolve aggregate microstructure following I2 staining. The automated image subtraction pipeline identified OM regions by the I Kedge induced brightness increases, achieving DSC values of 0.58–0.83 relative to an experienced annotator. Sensitivity analyses revealed that the reconstruction alpha value—optimized via the open-source tool TomocuPy—and the 3D registration slice count were the primary determinants of accuracy, providing a novel benchmark for dual-energy soil imaging. The pipeline without GPU acceleration achieved 9.6 to 43.2 times faster than manual annotation. Using GPU-accelerated image post-processing and affine transformation matrices, the pipeline successfully segmented OM elements for large-scale datasets (3232×3232 pixel, 2048 slices) within ~5200 s from raw file acquisition to segmented output. The high-throughput approach enables the quantification of OM spatial distribution across diverse and heterogeneous soil.

Soil microbial biomass↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

Automation-Accelerated Electrolyte Design Mitigates Solubility Competition between Redox-Active Molecules and Supporting Salts

In nonaqueous redox-flow batteries (NRFBs), redox-active organic molecules (ROMs) and supporting salts compete for solvation sites, limiting achievable energy density. We combine automated high-throughput experimentation (HTE) with camera-based saturation monitoring and quantitative NMR to measure paired (ROM, salt) solubilities across single and mixed organic solvents. Using 2,1,3-benzothiadiazole (BTZ) with lithium bis(trifluoromethanesulfonyl)imide (LiTFSI) as a model system, we find that a binary m-xylene/acetonitrile mixture dissolves ≈3 M of both BTZ and LiTFSI─surpassing the previously reported 2 M ceiling for neat acetonitrile─by leveraging complementary solvation (MX is BTZ-philic and salt-phobic; ACN stabilizes LiTFSI). A random-forest model (RMSE ≈ 0.24) trained on solvent descriptors highlights log P and salt concentration as dominant predictors and predicts MX/ACN ≈0.3/0.7 (v/v) to be near-optimal. These formulations retain practical viscosity and ∼5 mS·cm –1 conductivity at high loading. In conclusion, the workflow provides a reproducible, data-centric route to NRFB electrolyte design and motivates an open, standardized dual-solute solubility resource for accelerated electrolyte discovery.

Electrolytes↗

Automated Strain Construction for Biosynthetic Pathway Screening in Yeast

Automation accelerates the Design-Build-Test-Learn (DBTL) cycle for synthetic biology; however, most strain construction pipelines lack robotic integration. Here, in this study, we present the workflow design and source code for a modular, integrated protocol that automates the Build step in Saccharomyces cerevisiae. We programmed the Hamilton Microlab VANTAGE to integrate off-deck hardware via its central robotic arm, enabling automated steps that increased throughput to 2,000 transformations per week. We developed a user interface with the Hamilton VENUS software to support on-demand parameter customization. As a proof of concept, we screened a gene library in an engineered yeast strain producing verazine, a key intermediate in the biosynthesis of steroidal alkaloids. Our pipeline rapidly identified pathway bottlenecks and genes that enhanced verazine production by 2.0- to 5-fold. This technical note provides resources for synthetic biologists designing yeast workflows for biofoundries to screen libraries for pathway discovery/optimization, combinatorial biosynthesis, and protein engineering.

automation↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Segmentation method comparison for residual fiber length measurement across tiled microscopy images

Fiber length distribution (FLD), in part, governs mechanical properties in discontinuous fiber composites, yet manual measurement methods limit the high-throughput characterization needed for materials design optimization. This study compares deep learning segmentation approaches for automated FLD measurement in large-field microscopy, evaluating how method choice affects the microstructural descriptors used in structure-property-processing relationships. A critical challenge is that high-resolution microscopy images (10,000×10,000 pixels) must be tiled for deep learning analysis, fragmenting fibers at boundaries. We demonstrate that segmentation method proves crucial for measurement accuracy. For example, instance segmentation with Slicing Aided Hyper Inference (SAHI) preserves individual fiber integrity across tiles while semantic segmentation prioritizes speed. Comparing against manual measurement of extracted carbon fibers, YOLOv11-SAHI matched manual ground truth (238 μm weighted mean) with 40x speedup (4.5 vs 167 minutes per image). U-Net provides rapid quantification although it is at the cost of reduced accuracy due only reliably measuring stand-alone fibers. Our comparative analysis reveals that instance segmentation with SAHI better preserves length measurements while semantic segmentation prioritizes speed, providing empirical guidance for method selection. The characterization provides essential inputs for mechanical property prediction models and inverse design workflows, accelerating composite materials development cycles.

Additive manufacturing↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

MISPR : an open-source package for high-throughput multiscale molecular simulations

Computational tools provide a unique opportunity to study and design optimal materials by enhancing our ability to comprehend the connections between their atomistic structure and functional properties. However, designing materials with tailored functionalities is complicated due to the necessity to integrate various computational-chemistry software (not necessarily compatible with one another), the heterogeneous nature of the generated data, and the need to explore vast chemical and parameter spaces. The latter is especially important to avoid bias in scattered data points-based models and derive statistical trends only accessible by systematic datasets. Here, we introduce a robust high-throughput multi-scale computational infrastructure coined MISPR (Materials Informatics for Structure–Property Relationships) that seamlessly integrates classical molecular dynamics (MD) simulations with density functional theory (DFT). By enabling high-performance data analytics and coupling between different methods and scales, MISPR addresses critical challenges arising from the needs of automated workflow management and data provenance recording. The major features of MISPR include automated DFT and MD simulations, error handling, derivation of molecular and ensemble properties, and creation of output databases that organize results from individual calculations to enable reproducibility and transparency. In this work, we describe fully automated DFT workflows implemented in MISPR to compute various properties such as nuclear magnetic resonance chemical shift, binding energy, bond dissociation energy, and redox potential with support for multiple methods such as electron transfer and proton-coupled electron transfer reactions. The infrastructure also enables the characterization of large-scale ensemble properties by providing MD workflows that calculate a wide range of structural and dynamical properties in liquid solutions. MISPR employs the methodologies of materials informatics to facilitate understanding and prediction of phenomenological structure–property relationships, which are crucial to designing novel optimal materials for numerous scientific applications and engineering technologies.

36 MATERIALS SCIENCE↗

Automated segmentation of soft X-ray tomography: Native cellular structure with submicron resolution at high-throughput for whole-cell quantitative imaging in yeast

Soft X-ray tomography (SXT) is an invaluable tool for quantitatively analyzing cellular structures at suboptical isotropic resolution. However, it has traditionally depended on manual segmentation, limiting its scalability for large datasets. Here, we leverage a deep learning-based autosegmentation pipeline to segment and label cellular structures in hundreds of cells across three Saccharomyces cerevisiae strains. This task-based pipeline uses manual iterative refinement to improve segmentation accuracy for key structures, including the cell body, nucleus, vacuole, and lipid droplets, enabling high-throughput and precise phenotypic analysis. Using this approach, we quantitatively compared the three-dimensional (3D) whole-cell morphometric characteristics of wild-type, VPH1-GFP, and vac14 strains, uncovering detailed strain-specific cell and organelle size and shape variations. We show the utility of SXT data for precise 3D curvature analysis of entire organelles and cells and detection of fine morphological features using surface meshes. Our approach facilitates comparative analyses with high spatial precision and statistical throughput, uncovering subtle morphological features at the single-cell and population level. This workflow significantly enhances our ability to characterize cell anatomy and supports scalable studies on the mesoscale, with applications in investigating cellular architecture, organelle biology, and genetic research across diverse biological contexts.

Chen, Jianhua [Lawrence Berkeley National Laborato↗

Synthetic communities as a model for determining interactions between a biofertilizer chassis organism and native microbial consortia

Biofertilizers are critical for sustainable agriculture because they can replace ecologically disruptive chemical fertilizers while improving the trajectory of soil and plant health. However, for improving deployment, the persistence of biofertilizers within native soil consortia must be elucidated and enhanced. In this study we characterized a high-throughput, modular, and automation-friendly in vitro approach to screen for biofertilizer persistence within soil-derived consortia after co-cultivation with stable synthetic soil microbial communities (SynComs) obtained through a top-down cultivation process. Here, we profiled ~1200 SynComs isolated from various soil sources and cultivated in divergent media types, and we detected significant phylogenetic diversity (e.g. Shannon index >4) and richness (observed richness >400) across these communities. We observed high reproducibility in SynCom community structure from common soil and media types, which provided a testbed for assessing biofertilizer persistence within representative native consortia. Furthermore, we demonstrated that the screening method described herein can be coupled with microbial engineering to efficiently identify soil-derived SynComs in which an engineered biofertilizer organism (i.e. Bacillus subtilis) persists. Accordingly, we discovered that B. subtilis persisted in ~10% of SynComs that generally followed the diversity–invasion principle. Additionally, our approach enabled analysis of the ecological impact of B. subtilis inoculation on SynCom structure and profile alterations in community diversity and richness associated with the presence of a genetically modified model bacterium. Ultimately, this work has established a modular pipeline that could be integrated into a variety of microbiology/microbiome-relevant workflows or related applications that would benefit from assessment of the persistence of a specific organism of interest and its interaction with native consortia.

biofertilizers↗

High-Throughput Data Processing at FRIB Using ESnet

Real-time or nearly real-time (nearline) data processing methods are critical tools as detector technologies and data acquisition (DAQ) systems allow for higher data rates and volumes. The introduction of the energy sciences network (ESnet), a U.S. Department of Energy (DOE) supported high-speed network for scientific research, creates opportunities to leverage the computing power of DOE facilities like the National Energy Research Scientific Computing Center (NERSC). As a first step toward realizing a DOE Office of Science Integrated Research Infrastructure (IRI) pattern, an automated workflow was developed to remotely process data obtained from a nuclear physics experiment at the Facility for Rare Isotope Beams (FRIB) at NERSC with data transferred between FRIB and NERSC over ESnet. The workflow demonstrated the ability to process one week’s worth of experimental data in approximately 90 min and was used successfully for nearline analysis during a recently completed FRIB experiment. Here, a summary of the workflow development and results of recent demonstrations will be presented.

Data processing↗

A high-throughput experimentation platform for data-driven discovery in electrochemistry

Automating electrochemical analyses combined with artificial intelligence is poised to accelerate discoveries in renewable energy sciences and technologies. This study presents an automated high-throughput electrochemical characterization (AHTech) platform as a cost-effective and versatile tool for rapidly assessing liquid analytes. The Python-controlled platform combines a liquid handling robot, potentiostat, and customizable microelectrode bundles for diverse, reproducible electrochemical measurements in microtiter plates, minimizing chemical consumption and manual effort. To showcase the capability of AHTech, we screened a library of 180 small molecules as electrolyte additives for aqueous zinc metal batteries, generating data for training machine learning models to predict Coulombic efficiencies. Key molecular features governing additive performance were elucidated using Shapley Additive exPlanations and Spearman’s correlation, pinpointing high-performance candidates like cis-4-hydroxy-d-proline, which achieved an average Coulombic efficiency of 99.52% over 200 cycles. The workflow established herein is highly adaptable, offering a powerful framework for accelerating the exploration and optimization of extensive chemical spaces across diverse energy storage and conversion fields.

Lin, Dian-Zhao [Johns Hopkins University, Baltimor↗

DIVA/DeviceEditor v6.1.2

DIVA is an end-to-end DNA design and construction management platform that streamlines how researchers design, build, and receive sequence-verified DNA constructs. Through a web-based BioCAD interface (DeviceEditor), researchers independently design DNA constructs and submit them to a centralized queue with a single action. Designs progress transparently through standardized states which allow researchers to track status and access finished constructs via a central DNA repository. Submitted designs are reviewed by dedicated staff for feasibility and optimization, reducing costly failures and improving downstream execution. Automated DNA assembly software optimizes construction strategies by reusing existing parts where possible and sourcing synthetic DNA only when needed. Standardized, sequence-agnostic assembly methods enable many independent constructs to be built in parallel using lab automation, dramatically increasing throughput. High-throughput next-generation sequencing is used to verify construct accuracy, with flexible platforms selected based on task requirements. Throughout the process, detailed success and failure data are captured and analyzed, enabling continuous improvement of assembly protocols. Compared to traditional, manual DNA construction workflows, DIVA offers higher scalability, transparency, reproducibility, and data-driven optimization.

Plahar, Hector [Lawrence Berkeley National Laborat↗

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L↗