Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Traditional Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Interpretable boosted-decision-tree analysis for the Majorana Demonstrator

The Majorana Demonstrator is a leading experiment searching for neutrinoless double-beta decay with high purity germanium detectors (HPGe). Machine learning provides a new way to maximize the amount of information provided by these detectors, but the data-driven nature makes it less interpretable compared to traditional analysis. An interpretability study reveals the machine's decision-making logic, allowing us to learn from the machine to feedback to the traditional analysis. In this work, we have presented the first machine learning analysis of the data from the Majorana Demonstrator; this is also the first interpretable machine learning analysis of any germanium detector experiment. Two gradient boosted decision tree models are trained to learn from the data, and a game-theory-based model interpretability study is conducted to understand the origin of the classification power. By learning from data, this analysis recognizes the correlations among reconstruction parameters to further enhance the background rejection performance. By learning from the machine, this analysis reveals the importance of new background categories to reciprocally benefit the standard Majorana analysis. This model is highly compatible with next-generation germanium detector experiments like LEGEND since it can be simultaneously trained on a large number of detectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accurate and Rapid Forecasts for Geologic Carbon Storage via Learning-Based Inversion-Free Prediction

Carbon capture and storage (CCS) is one approach being studied by the U.S. Department of Energy to help mitigate global warming. The process involves capturing CO 2 emissions from industrial sources and permanently storing them in deep geologic formations (storage reservoirs). However, CCS projects generally target “green field sites,” where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir. Consequently, ensemble-based approaches are often used to forecast multiple realizations prior to CO 2 injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Thus, a system for generating accurate, timely forecasts of pressure buildup and CO 2 movement and distribution within the storage reservoir and for updating those forecasts via monitoring measurements becomes crucial. This study proposes a learning-based prediction method that can accurately and rapidly forecast spatial distribution of CO 2 concentration and pressure with uncertainty quantification without relying on traditional inverse modeling. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO 2 storage site operators with an effective tool for timely and informative decision making based on limited simulation and monitoring data.

58 GEOSCIENCES↗

Development and Evaluation of a General Drag Model for Gas-Solid Flows via Deep Learning

This project presents the development and evaluation of a general drag model for gas–solid multiphase flows using deep learning techniques. A comprehensive database of more than 4,000 experimental and numerical data points for spherical and non spherical particles was compiled, incorporating geometric features such as sphericity, aspect ratio, and orientation. Several predictive approaches—including traditional em pirical correlations, machine learning, and deep neural networks—were benchmarked, with the proposed Drag Coefficient Correlation-aided Deep Neural Network (DCC DNN) demonstrating superior accuracy. To account for particle–particle interactions, additional drag data were generated using CFD-based simulations of packed and flu idized beds, leading to the development of a retrained model capable of incorporat ing volume fraction effects. Integration of the trained model with the MFiX CFD solver was achieved using FTorch, enabling drag predictions during discrete element method (DEM) simulations. Validation against experimental data for single particles and fluidized beds confirmed the model’s improved predictive ability, particularly for non-spherical geometries. While the model performed strongly under fluidized con ditions, limitations remained in unfluidized regimes, suggesting a need for expanded datasets. Overall, this study demonstrates the feasibility of combining deep learning with physics-informed CFD to improve drag modeling for gas–solid flows, with promis ing implications for scaling multiphase simulations in industrial applications.

42 ENGINEERING↗

Machine-learning-based spectral methods for partial differential equations

Spectral methods are an important part of scientific computing’s arsenal for solving partial differential equations (PDEs). However, their applicability and effectiveness depend crucially on the choice of basis functions used to expand the solution of a PDE. The last decade has seen the emergence of deep learning as a strong contender in providing efficient representations of complex functions. In the current work, we present an approach for combining deep neural networks with spectral methods to solve PDEs. In particular, we use a deep learning technique known as the Deep Operator Network (DeepONet) to identify candidate functions on which to expand the solution of PDEs. We have devised an approach that uses the candidate functions provided by the DeepONet as a starting point to construct a set of functions that have the following properties: (1) they constitute a basis, (2) they are orthonormal, and (3) they are hierarchical, i.e., akin to Fourier series or orthogonal polynomials. We have exploited the favorable properties of our custom-made basis functions to both study their approximation capability and use them to expand the solution of linear and nonlinear time-dependent PDEs. The proposed approach advances the state of the art and versatility of spectral methods and, more generally, promotes the synergy between traditional scientific computing and machine learning.

97 MATHEMATICS AND COMPUTING↗

Evaluation of the economic implications of varied pressure drawdown strategies generated using a real-time, rapid predictive, multi-fidelity model for unconventional oil and gas wells

Experience has suggested that pressure maintenance in hydraulically fractured reservoirs via lower, more sustained production drawdowns may offer improved cumulative recovery and overall resource extraction efficiency compared to more rapid drawdown approaches aimed at generating high initial production. However, given the inherent variability of oil and natural gas markets, operators pursue production strategies that maximize profitability over resource extraction efficiency. This study focuses on evaluating the implications of contrasting pressure drawdown strategies on the long-term production and resulting economics for a real, producing unconventional gas well in the Marcellus Shale of the Appalachian Basin using a techno-economic analysis approach. Our research combines elements of well-specific horizontal well design, production forecasting, equipment sizing and capital cost estimation, operating cost estimation, and revenue and tax calculations. Gas production forecast outlook scenarios were generated under varying pressure drawdowns using two approaches: 1) a novel physics-informed machine learning workflow and 2) traditional reservoir simulation. A discounted cash flow model was used to evaluate the resulting economic implications for each drawdown scenario—generating output for exploring the coupled effect of factors like the timing and volume of gas production, prevailing economic and market conditions for natural gas, and overall estimated ultimate recovery on profitability metrics such as internal rate of return and net present value. Results show that there is potential to maximize the cumulative gas produced in the specific case study well by employing a lower pressure drawdown. Conversely, the greatest profitability is achieved using rapid drawdown as signified by a small, specific subset of our outlook scenarios. On an averaging basis, we find that the combinations of highest cumulative producing and most profitable scenarios occur under lower drawdowns with long (>40 years) producing timeframes, but require higher relative gas price and lower discounting considerations. Further, the machine learning predictive outlooking capability proved effective for enabling rapid generation of a multitude of scenario forecasts. As a result, a variety of prominent example cases could be generated to strike the balance of greater productivity and economic return given their associated producing features and economic conditions when compared to similar producing scenarios—critical insight that offers improved decision support for unconventional oil and gas operations.

42 ENGINEERING↗

Large-scale scenarios of electric vehicle charging with a data-driven model of control

Transportation electrification is forecast to bring millions of new electric vehicles to roads worldwide this decade. Planning to support those vehicles depends on detailed scenarios of their electricity demand in both uncontrolled and controlled or smart charging scenarios. In this work, we present a novel modeling approach to enable rapid generation of demand estimates that represent the impact of controlled charging for large-scale scenarios with millions of individual drivers. To model the effect of load modulation control on aggregate charging profiles, we propose a novel machine learning approach that replaces traditional optimization approaches. We demonstrate its performance modeling workplace charging control under a range of electricity rate schedules, achieving small errors (2.5%–4.5%) while accelerating computations by more than 4000 times. To generate the uncontrolled charging demand for scenarios with residential, workplace, and public charging we use statistical representations of a large data set of real charging sessions. We demonstrate the methodology by generating diverse sets of scenarios for California's charging demand in 2030 which consider multiple charging segments and controls, each run locally in under 50 s. We further demonstrate support for rate design by modeling the large-scale impact of a new, custom rate schedule for workplace charging.

33 ADVANCED PROPULSION SYSTEMS↗

Identification of new marker genes from plant single‐cell RNA‐seq data using interpretable machine learning methods

Summary An essential step in the analysis of single‐cell RNA sequencing data is to classify cells into specific cell types using marker genes. In this study, we have developed a machine learning pipeline called single‐cell predictive marker (SPmarker) to identify novel cell‐type marker genes in the Arabidopsis root. Unlike traditional approaches, our method uses interpretable machine learning models to select marker genes. We have demonstrated that our method can: assign cell types based on cells that were labelled using published methods; project cell types identified by trajectory analysis from one data set to other data sets; and assign cell types based on internal GFP markers. Using SPmarker, we have identified hundreds of new marker genes that were not identified before. As compared to known marker genes, the new marker genes have more orthologous genes identifiable in the corresponding rice single‐cell clusters. The new root hair marker genes also include 172 genes with orthologs expressed in root hair cells in five non‐ Arabidopsis species, which expands the number of marker genes for this cell type by 35–154%. Our results represent a new approach to identifying cell‐type marker genes from scRNA‐seq data and pave the way for cross‐species mapping of scRNA‐seq data in plants.

54 ENVIRONMENTAL SCIENCES↗

Optimal sensor placement for reconstructing wind pressure field around buildings using compressed sensing

Deciding how to optimally deploy sensors in a large, complex, and spatially extended structure is critical to ensure that the surface pressure field is accurately captured for subsequent analysis and design. In some cases, reconstruction of missing data is required in downstream tasks such as the development of digital twins. Here, this paper presents a data-driven sparse sensor selection algorithm, aiming to provide the most information contents for reconstructing aerodynamic characteristics of wind pressures over tall building structures parsimoniously. The algorithm first fits a set of basis functions to the training data, then applies a computationally efficient QR algorithm that ranks existing pressure sensors in order of importance based on the state reconstruction to this tailored basis. The findings of this study show that the proposed algorithm successfully re- constructs the aerodynamic characteristics of tall buildings from sparse measurement locations, generating stable and optimal solutions across a range of conditions. As a result, this study serves as a promising first step toward leveraging the success of data-driven and machine learning algorithms to supplement traditional genetic algorithms currently used in wind engineering.

42 ENGINEERING↗

Discovering nuclear models from symbolic machine learning

Numerous phenomenological nuclear models have been proposed to describe specific observables within different regions of the nuclear chart. However, developing a unified model that describes the complex behavior of all nuclei remains an open challenge. Here, we explore whether symbolic Machine Learning (ML) can rediscover traditional nuclear physics models or identify alternatives with improved simplicity, fidelity, and predictive power. To address this challenge, we developed a Multi-objective Iterated Symbolic Regression approach that handles symbolic regressions over multiple target observables, accounts for experimental uncertainties and is robust against high-dimensional problems. As a proof of principle, we applied this method to describe the nuclear binding energies and charge radii of light and medium mass nuclei. Our approach identified simple analytical relationships based on the number of protons and neutrons, providing interpretable models with precision comparable to state-of-the-art nuclear models. Additionally, we integrated this ML-discovered model with an existing complementary model to estimate the limits of nuclear stability. These results highlight the potential of symbolic ML to develop accurate nuclear models and guide our description of complex many-body problems.

Nuclear structure↗

Modeling of the metal–insulator transition temperature in alio-valently doped VO 2 through symbolic regression

The correlated semiconductor vanadium dioxide (VO 2 ) exhibits an insulator–metal transition (IMT) near room temperature, which is of interest in various device applications. Precise IMT temperature control is crucial to determine the use cases across technologies such as thermochromic windows, actuators for robots or neuronal oscillators. Doping the cation or anion sites can modulate the IMT by several tens of degrees and control hysteresis. However, modeling the effects of control parameters (e.g., doping concentration, type of dopants) is challenging due to complex experimental procedures and limited data, hindering the use of traditional data-driven machine learning approaches. Symbolic regression (SR) can bridge this gap by identifying nonlinear expressions connecting key input parameters to target properties, even with small data sets. In this work, we develop SR models to capture the IMT trends in VO 2 influenced by different dopant parameters. Using experimental data from the literature, our study reveals a dual nature of the IMT temperature with varying tungsten (W) doping concentrations. The symbolic model captures data trends and accounts for experimental variability, providing a complementary approach to first-principles calculations. Our feature-driven analysis across a broader class of dopants informs selectivity and provides qualitative insights into tuning phase transition properties valuable for neuromorphic computing and thermochromic windows.

36 MATERIALS SCIENCE↗

Effective cosmic density field reconstruction with convolutional neural network

ABSTRACT We present a cosmic density field reconstruction method that augments the traditional reconstruction algorithms with a convolutional neural network (CNN). Following previous work, the key component of our method is to use the reconstructed density field as the input to the neural network. We extend this previous work by exploring how the performance of these reconstruction ideas depends on the input reconstruction algorithm, the reconstruction parameters, and the shot noise of the density field, as well as the robustness of the method. We build an eight-layer CNN and train the network with reconstructed density fields computed from the Quijote suite of simulations. The reconstructed density fields are generated by both the standard algorithm and a new iterative algorithm. In real space at z = 0, we find that the reconstructed field is 90 per cent correlated with the true initial density out to $k\sim 0.5 \, \mathrm{ h}\, \rm {Mpc}^{-1}$, a significant improvement over $k\sim 0.2 \, \mathrm{ h}\, \rm {Mpc}^{-1}$ achieved by the input reconstruction algorithms. We find similar improvements in redshift space, including an improved removal of redshift space distortions at small scales. We also find that the method is robust across changes in cosmology. Additionally, the CNN removes much of the variance from the choice of different reconstruction algorithms and reconstruction parameters. However, the effectiveness decreases with increasing shot noise, suggesting that such an approach is best suited to high density samples. This work highlights the additional information in the density field beyond linear scales as well as the power of complementing traditional analysis approaches with machine learning techniques.

Astronomy & Astrophysics↗

A globally sampled high-resolution hand-labeled validation dataset for evaluating surface water extent maps

Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.

54 ENVIRONMENTAL SCIENCES↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗

Stable Machine‐Learning Parameterization of Subgrid Processes in a Comprehensive Atmospheric Model Learned From Embedded Convection‐Permitting Simulations

Modern climate projections often suffer from inadequate spatial and temporal resolution due to computational limitations, resulting in inaccurate representations of sub-grid processes. A promising technique to address this is the multiscale modeling framework (MMF), which embeds a kilometer-resolution cloud-resolving model (CRM) within each atmospheric column of a host climate model to replace traditional convection and cloud parameterizations. Machine learning offers a unique opportunity to make MMF more accessible by emulating the embedded CRM and reducing its substantial computational cost. Although many studies have demonstrated proof-of-concept success of achieving stable hybrid simulations, it remains a challenge to achieve near operational-level success with real geography and comprehensive variable emulation that includes, for example, explicit cloud condensate coupling. In this study, we present a stable hybrid model capable of integrating for at least 5 years with near operational-level complexity, including coarse-grid geography, seasonality, explicit cloud condensate and wind predictions, and land coupling. Our model demonstrates skillful online performance, achieving a 5-year zonal mean tropospheric temperature bias within 2 K, water vapor bias within 1 g/kg, and a precipitation root mean square error of 0.96 mm/day. Key factors contributing to our online performance include an expressive U-Net architecture and physical thermodynamic constraints for microphysics. With microphysical constraints mitigating unrealistic cloud formation, our work is the first to demonstrate realistic multi-year cloud condensate climatology under the MMF framework. Despite these advances, online diagnostics reveal persistent biases in certain regions, highlighting the need for innovative strategies to further optimize online performance.

Hu, Zeyuan [NVIDIA Corporation, Santa Clara, CA (U↗

A systematic review of machine learning in groundwater monitoring

With increasing concerns about water scarcity, groundwater has become crucial since this resource provides most of the freshwater needs. However, various human and natural activities often contaminate the groundwater, making it unsuitable for use. Over the years, scientists and engineers have used many methods to predict and track groundwater contamination as part of environmental monitoring. Consequently, there is an urgent need for improved methods, particularly in the face of increasing contamination. Machine learning has sometimes been used to monitor groundwater, air quality, and climate. Traditional methods must be improved due to the complexity and large amount of environmental data. This includes using hybrid models that combine traditional and new techniques. Despite the use of machine learning in many scientific areas, there is a lack of comprehensive reviews focusing on its use in environmental monitoring, especially groundwater monitoring. We aim to fill this gap by exploring machine-learning applications in groundwater monitoring. We discuss relevant methods, their limitations, and future potential. We summarize research on automating data processing and model training using groundwater sensor data. Our research underscores the transformative potential of machine learning to revolutionize long-term groundwater monitoring and contamination detection, providing valuable insights for future research and practical applications.

AI/ML↗

Characterizing Machine Learning I/O Workloads on Leadership Scale HPC Systems

High performance computing (HPC) is no longer solely limited to traditional workloads such as simulation and modeling. With the increase in the popularity of machine learning (ML) and deep learning (DL) technologies, we are observing that an increasing number of HPC users are incorporating ML methods into their workflow and scientific discovery processes, across a wide spectrum of science domains such as biology, earth science, and physics. This gives rise to a diverse set of I/O patterns than the traditional checkpoint/restart-based HPC I/O behavior. The details of the I/O characteristics of such ML I/O workloads have not been studied extensively for large-scale leadership HPC systems. This paper aims to fill that gap by providing an in-depth analysis to gain an understanding of the I/O behavior of ML I/O workloads using darshan - an I/O characterization tool designed for lightweight tracing and profiling. We study the darshan logs of more than 23, 000 HPC ML I/O jobs over a time period of one year running on Summit - the second-fastest supercomputer in the world. This paper provides a systematic I/O characterization of ML I/O jobs running on a leadership scale supercomputer to understand how the I/O behavior differs across science domains and the scale of workloads, and analyze the usage of parallel file system and burst buffer by ML I/O workloads.

Paul, Arnab↗

Machine Learning-Enabled Image Classification for Automated Electron Microscopy

Abstract Traditionally, materials discovery has been driven more by evidence and intuition than by systematic design. However, the advent of “big data” and an exponential increase in computational power have reshaped the landscape. Today, we use simulations, artificial intelligence (AI), and machine learning (ML) to predict materials characteristics, which dramatically accelerates the discovery of novel materials. For instance, combinatorial megalibraries, where millions of distinct nanoparticles are created on a single chip, have spurred the need for automated characterization tools. This paper presents an ML model specifically developed to perform real-time binary classification of grayscale high-angle annular dark-field images of nanoparticles sourced from these megalibraries. Given the high costs associated with downstream processing errors, a primary requirement for our model was to minimize false positives while maintaining efficacy on unseen images. We elaborate on the computational challenges and our solutions, including managing memory constraints, optimizing training time, and utilizing Neural Architecture Search tools. The final model outperformed our expectations, achieving over 95% precision and a weighted F-score of more than 90% on our test data set. This paper discusses the development, challenges, and successful outcomes of this significant advancement in the application of AI and ML to materials discovery.

Materials Science↗

Computationally Accelerated Discovery and Experimental Demonstration of High-Performance Materials for Advanced Solar Thermochemical Hydrogen Production

This project achieved its overarching goal of accelerating the discovery and validation of solar thermochemical hydrogen (STCH) materials through a tightly integrated approach that combined high-throughput computational screening, advanced machine learning (ML), and experimental testing. Guided by the objectives outlined in the Statement of Project Objectives (SOPO), our work fulfilled all major milestones across four technical tasks and delivered scientific breakthroughs and practical tools that significantly exceeded the original scope of the project. We began by addressing the challenge of predicting material phase stability through machine learning. A novel Python module was developed to generate thousands of meaningful features from composition, structure, and electronic properties, enabling rapid and reproducible ML model development. Using these tools, we trained a model to predict temperature-dependent Gibbs energies (G(T)) for inorganic crystalline materials with near-chemical accuracy—roughly 40 meV/atom—marking the first such descriptor of its kind. We also introduced a new machine-learned tolerance factor, τ, that accurately predicted perovskite formability with over 90% success, outperforming traditional heuristic models, such as the Goldschmidt tolerance factor. These capabilities allowed for rapid and accurate predictions of phase stability across a vast oxide composition space, setting the stage for high-throughput thermodynamic screening. Building on this foundation, we conducted an extensive computational screening of candidate STCH oxide materials. Over 1.1 million perovskite compositions were evaluated using the τ descriptor, leading to the identification of more than 27,000 predicted stable structures. Using density functional theory (DFT), we refined over 68,000 multinary perovskite structures and computed oxygen vacancy formation energies for over 1,300 ternary and double perovskites. These calculations enabled us to isolate compounds with redox behavior consistent with STCH requirements and resulted in a public dataset now hosted on the Materials Project. Recognizing that thermodynamic screening alone is insufficient, we addressed kinetic limitations by developing a suite of tools to estimate transition state (TS) energies for key redox reactions. We implemented a novel bounding approach that provides lower and upper estimates of TS energies with dramatically reduced computational cost, requiring less than 10% of the CPU time of a full nudged elastic band (NEB) calculation while maintaining high accuracy. This enabled rapid evaluation of over 200 reaction pathways across 90 materials. To further accelerate screening, we developed a SISSO-based ML model to predict diffusion barriers with a 96.7% success rate in classifying fast vs. slow materials, supporting a robust, data-driven framework for assessing redox kinetics. Experimental validation was critical to confirming the predictive power of our models. We synthesized and tested a wide array of candidate materials, including Mn-doped hercynite and several Gd- and La-based perovskites. Notably, Sr 0.4 Gd 0.6 Mn 0.6 Al 0.4 O 3 (SGMA) and Gd 0.5 La 0.5 Co 0.5 Fe 0.5 O 3 (GLCF) emerged as leading STCH materials, exhibiting robust redox cycling and high hydrogen yields exceeding 150 µmol H 2 /g per cycle. These materials also retained over 50% of their hydrogen productivity under high-conversion conditions (H 2 O:H 2 = 1333:1), demonstrating strong thermodynamic favorability and promising performance under industrially relevant scenarios. Additional candidates, such as La 2 MnNiO 6 (L2MN), were found to produce even higher yields than ceria under standard STCH conditions. Our collaborators at Sandia National Laboratories confirmed these findings using high-temperature X-ray diffraction and thermogravimetric analysis, observing stable phase evolution and reversible redox activity. In several respects, the project went beyond the goals initially outlined in the SOPO. We published 17 peer-reviewed articles, including a large dataset of over 66,000 theoretical perovskites and a new structure prediction method (SPuDS-DFT) that accurately identifies ground-state structures at a fraction of the cost of traditional DFT. We demonstrated that our machine-learned G(T) model offers accuracy rivaling quasiharmonic calculations while being orders of magnitude faster. In partnership with the Materials Project, we made our datasets openly available, providing a powerful new resource for the broader materials science community. The combined computational and experimental advances of this project represent a significant advance in STCH materials discovery. By creating a robust, generalizable, and open workflow for thermodynamic and kinetic screening, and validating key findings through synthesis and reactor testing, we have provided a practical and scalable pathway for the rapid identification of new redox-active materials. The tools, data, and materials developed under this project are already supporting ongoing research and have laid the groundwork for the next generation of solar fuel technologies.

08 HYDROGEN↗