Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

An Alternative Ensemble Streamflow Prediction Approach Using Improved Subseasonal Precipitation Forecasts from the North America Multi-Model Ensemble Phase II

In this article, streamflow forecasting at a subseasonal time scale (10–30 days into the future) is important for various human activities. The ensemble streamflow prediction (ESP) is a widely applied technique for subseasonal streamflow forecasting. However, ESP’s reliance on the randomly resampled historical precipitation limits its predictive capability. Available dynamical subseasonal precipitation forecasts provide an alternative to the randomly resampled precipitation in ESP. Prior studies found the predictive performance of raw subseasonal precipitation forecast is limited in many regions such as the central south of the United States, which raises questions about its effectiveness in assisting streamflow forecasting. To further assess the hydrologic applicability of dynamical subseasonal precipitation forecasts, we test the subseasonal precipitation forecast from North America Multi-Model Ensemble Phase II (NMME-2) at four watersheds in the central south region of the United States. The subseasonal precipitation forecasts are postprocessed with bias correction and spatial disaggregation (BCSD) to correct bias and improve spatial resolution before replacing the randomly resampled precipitation in ESP for streamflow predictions. The performance of the resulting streamflow predictions is benchmarked with ESP. Evaluation is conducted using Kling–Gupta Efficiency (KGE), continuous ranked probability score (CRPS), probability of detection (POD), false alarm ratios (FARs), as well as reliability diagrams. Our results suggest that BCSD-corrected subseasonal precipitation forecasts lead to overall improved streamflow predictions due to added skills in winter and spring. Our results also suggest that BCSD-corrected subseasonal precipitation forecasts lead to improved predictions on the occurrence of high-percentile streamflow values above 75%. Overall, BCSD-corrected subseasonal precipitation has shown promising performance, highlighting its potential broader applications for river and flood forecasting.

54 ENVIRONMENTAL SCIENCES↗

Statistical upscaling of ecosystem CO 2 fluxes across the terrestrial tundra and boreal domain: Regional patterns and uncertainties

Abstract The regional variability in tundra and boreal carbon dioxide (CO 2 ) fluxes can be high, complicating efforts to quantify sink‐source patterns across the entire region. Statistical models are increasingly used to predict (i.e., upscale) CO 2 fluxes across large spatial domains, but the reliability of different modeling techniques, each with different specifications and assumptions, has not been assessed in detail. Here, we compile eddy covariance and chamber measurements of annual and growing season CO 2 fluxes of gross primary productivity (GPP), ecosystem respiration (ER), and net ecosystem exchange (NEE) during 1990–2015 from 148 terrestrial high‐latitude (i.e., tundra and boreal) sites to analyze the spatial patterns and drivers of CO 2 fluxes and test the accuracy and uncertainty of different statistical models. CO 2 fluxes were upscaled at relatively high spatial resolution (1 km 2 ) across the high‐latitude region using five commonly used statistical models and their ensemble, that is, the median of all five models, using climatic, vegetation, and soil predictors. We found the performance of machine learning and ensemble predictions to outperform traditional regression methods. We also found the predictive performance of NEE‐focused models to be low, relative to models predicting GPP and ER. Our data compilation and ensemble predictions showed that CO 2 sink strength was larger in the boreal biome (observed and predicted average annual NEE −46 and −29 g C m −2 yr −1 , respectively) compared to tundra (average annual NEE +10 and −2 g C m −2 yr −1 ). This pattern was associated with large spatial variability, reflecting local heterogeneity in soil organic carbon stocks, climate, and vegetation productivity. The terrestrial ecosystem CO 2 budget, estimated using the annual NEE ensemble prediction, suggests the high‐latitude region was on average an annual CO 2 sink during 1990–2015, although uncertainty remains high.

Virkkala, Anna‐Maria↗

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

The Seasonal-to-Multiyear Large Ensemble (SMYLE) prediction system using the Community Earth System Model version 2

Abstract. The potential for multiyear prediction of impactful Earth system change remains relatively underexplored compared to shorter (subseasonal to seasonal) and longer (decadal) timescales. In this study, we introduce a new initialized prediction system using the Community Earth System Model version 2 (CESM2) that is specifically designed to probe potential and actual prediction skill at lead times ranging from 1 month out to 2 years. The Seasonal-to-Multiyear Large Ensemble (SMYLE) consists of a collection of 2-year-long hindcast simulations, with four initializations per year from 1970 to 2019 and an ensemble size of 20. A full suite of output is available for exploring near-term predictability of all Earth system components represented in CESM2. We show that SMYLE skill for El Niño–Southern Oscillation is competitive with other prominent seasonal prediction systems, with correlations exceeding 0.5 beyond a lead time of 12 months. A broad overview of prediction skill reveals varying degrees of potential for useful multiyear predictions of seasonal anomalies in the atmosphere, ocean, land, and sea ice. The SMYLE dataset, experimental design, model, initial conditions, and associated analysis tools are all publicly available, providing a foundation for research on multiyear prediction of environmental change by the wider community.

54 ENVIRONMENTAL SCIENCES↗

Revealing the role of redox reaction selectivity and mass transfer in current–voltage predictions for ensembles of photocatalysts

Photocatalysts are conceptually simple reaction units where nanoscale semiconductors integrated with catalysts drive a pair of redox reactions on illumination. However, the proximity of reaction sites performing cathodic and anodic reactions poses dire challenges to realize large light-to-fuel conversion efficiencies. In this study, a powerful, yet straightforward, equivalent-circuit detail-balance modeling framework is developed and applied to evaluate the performance of photocatalytic systems featuring multiple light absorbers. Specifically, low bandgap iridium-doped strontium titanate is modeled as a Z-scheme photocatalyst to achieve desirable hydrogen evolution and iron-based redox shuttle oxidation reactions. Our model has unique capabilities to simulate competing redox reactions and address mass-transfer limitations. In a significant departure from state-of-the-art circuit models, our study develops tools to perform load-line analyses by incorporating a net electrochemical load curve that includes both desired and competing redox reactions. Consequently, reaction selectivity is predicted from equivalent circuit models for photocatalytic and photoelectrochemical systems. Our investigation into ensembles comprised of multiple, semi-transparent light absorbers reveals their potential to outperform a single, optically thick light absorber, particularly when operated under mass-transfer-limited conditions. However, this outcome hinges on minimizing mass-transfer rates of select redox species to prevent undesired reactions of hydrogen oxidation and/or redox shuttle reduction. Our findings demonstrate that reaction selectivity can be achieved by tuning asymmetry in redox species mass-transfer even with perfectly symmetric electrocatalytic charge-transfer coefficients. The influences of various kinetic, mass-transfer, and thermodynamic parameters are explored to offer crucial insights for synthesis of the next-generation of photocatalysts and selective coatings, and reactor designs.

25 ENERGY STORAGE↗

Ensemble Learning, Prediction and Li-Ion Cell Charging Cycle Divergence

In recent years, the pervasive use of lithium ion (Li-ion) batteries in applications such as cell phones, laptop computers, electric vehicles, and grid energy storage systems has prompted the development of specialized battery management systems (BMS). The primary goal of a BMS is to maintain a reliable and safe battery power source while maximizing the calendar life and performance of the cells. To maintain safe operation, a BMS should be programmed to minimize degradation and prevent damage to a Li-ion cell, which can lead to thermal runaway. Cell damage can occur over time if a BMS is not properly configured to avoid overcharging and discharging. To prevent cell damage, efficient and accurate cell charging cycle characteristics algorithms must be employed. In this paper, computationally efficient and accurate ensemble learning algorithms capable of detecting Li-ion cell charging irregularities are described. Additionally, it is shown using machine and deep learning that it is possible to accurately and efficiently detect when a cell has experienced thermal and electrical stress due to cell overcharging by measuring charging cycle divergence.

25 ENERGY STORAGE↗

Assessments of epistemic uncertainty using Gaussian stochastic weight averaging for fluid-flow regression

Here, we use Gaussian stochastic weight averaging (SWAG) to assess the epistemic uncertainty associated with neural-network-based function approximation relevant to fluid flows. SWAG approximates a posterior Gaussian distribution of each weight, given training data, and a constant learning rate. Having access to this distribution, it is able to create multiple models with various combinations of sampled weights, which can be used to obtain ensemble predictions. The average of such an ensemble can be regarded as the 'mean estimation', whereas its standard deviation can be used to construct 'confidence intervals', which enable us to perform uncertainty quantification (UQ) with regard to the training process of neural networks. We utilize representative neural-network-based function approximation tasks for the following cases: (i) a two-dimensional circular-cylinder wake; (ii) the DayMET dataset (maximum daily temperature in North America); (iii) a three-dimensional square-cylinder wake; and (iv) urban flow, to assess the generalizability of the present idea for a wide range of complex datasets. SWAG-based UQ can be applied regardless of the network architecture, and therefore, we demonstrate the applicability of the method for two types of neural networks: (i) global field reconstruction from sparse sensors by combining convolutional neural network (CNN) and multi-layer perceptron (MLP); and (ii) far-field state estimation from sectional data with two-dimensional CNN. We find that SWAG can obtain physically-interpretable confidence-interval estimates from the perspective of epistemic uncertainty. This capability supports its use for a wide range of problems in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Probabilistic Cloud Optimized Day-Ahead Forecasting System Based on WRF-Solar (Final Report)

The most persistent challenge in both intraday and day-ahead solar forecasting is to get numerical weather prediction models to produce the right type of clouds with the right frequency at the right time and place. Another challenge is to understand and communicate the forecast uncertainty. The objective of this project was to develop an optimized ensemble-based solar irradiance forecasting system that will (1) demonstrably improve the current state-of-the-art solar forecasts from the deterministic Weather Research and Forecasting-Solar (WRF-Solar) model and (2) provide probabilistic forecasts for grid operations. This probabilistic solar forecasting system, referred to as the WRF-Solar Ensemble Prediction System (WRF-Solar EPS), aims to significantly enhance both the intraday and the day-ahead solar forecasting capability for grid operations. This technical report summarizes the work performed in the past 3 years through a collaboration between the National Renewable Energy Laboratory and the National Center for Atmospheric Research as part of the U.S. Department of Energy's Solar Forecasting 2 program that aims to improve the accuracy of solar energy forecasts and enable increased deployment of solar energy on the electric grid.

14 SOLAR ENERGY↗

The Impact of Stochastic Perturbations in Physics Variables for Predicting Surface Solar Irradiance

We present a probabilistic framework tailored for solar energy applications referred to as the Weather Research and Forecasting-Solar ensemble prediction system (WRF-Solar EPS). WRF-Solar EPS has been developed by introducing stochastic perturbations into the most relevant physical variables for solar irradiance predictions. In this study, we comprehensively discuss the impact of the stochastic perturbations of WRF-Solar EPS on solar irradiance forecasting compared to a deterministic WRF-Solar prediction (WRF-Solar DET), a stochastic ensemble using the stochastic kinetic energy backscatter scheme (SKEBS), and a WRF-Solar multi-physics ensemble (WRF-Solar PHYS). The performances of the four forecasts are evaluated using irradiance retrievals from the National Solar Radiation Database (NSRDB) over the contiguous United States. We focus on the predictability of the day-ahead solar irradiance forecasts during the year of 2018. The results show that the ensemble forecasts improve the quality of the forecasts, compared to the deterministic prediction system, by accounting for the uncertainty derived by the ensemble members. However, the three ensemble systems are under-dispersive, producing unreliable and overconfident forecasts due to a lack of calibration. In particular, WRF-Solar EPS produces less optically thick clouds than the other forecasts, which explains the larger positive bias in WRF-Solar EPS (31.7 W/m 2 ) than in the other models (22.7–23.6 W/m 2 ). This study confirms that the WRF-Solar EPS reduced the forecast error by 7.5% in terms of the mean absolute error (MAE) compared to WRF-Solar DET, and provides in-depth comparisons of forecast abilities with the conventional scientific probabilistic approaches (i.e., SKEBS and a multi-physics ensemble). Guidelines for improving the performance of WRF-Solar EPS in the future are provided.

14 SOLAR ENERGY↗

Evaluating WRF-Solar EPS cloud mask forecast using the NSRDB

Improving the accuracy of day-ahead solar forecasts using numerical weather prediction models requires improving the forecasting of cloud occurrence and properties. Validating cloud forecasts is challenging because this requires evaluating different types of clouds over a wide variety of regions. This study analyzes the cloud occurrence (or cloud mask) over the contiguous United States (CONUS), predicted by the Weather Research and Forecasting-Solar Ensemble Prediction System (WRF-Solar EPS), to identify the strengths and limitations of the model in reproducing cloud fields. To enable the in-depth analysis of cloud mask forecasts covering CONUS, we use satellite observations from the National Solar Radiation Database (NSRDB). Two evaluation methods are implemented to consider all clouds and partially removed clouds from the 2-km NSRDB in evaluating the 9-km WRF-Solar EPS cloud mask. Cloud detection metrics as well as the frequency of cloud occurrence are used to quantify the monthly performance of WRF-Solar EPS. Mismatched cloud frequency (MCF) is used to assess the model's capability to predict different types of clouds, which are classified using three levels of cloud optical depth (COD) and cloud top height (CTH). The day-ahead forecasts covering the full year of 2018 demonstrate that WRF-Solar EPS produces MCFs ranging from 27%-46%, 13%-34%, and 8%-19% for thin, mid-thickness, and thick clouds, respectively. For three CTH levels, the model shows MCFs ranging from 19%-46%, 16%-33%, and 8%-27% for low-level, middle-level, and high-level clouds, respectively. This comprehensive characterization of model performance helps identify model weakness and will eventually lead to improvements in cloud and solar radiation forecasting.

14 SOLAR ENERGY↗

Machine learning for postprocessing ensemble streamflow forecasts

Skillful streamflow forecasts can inform decisions in various areas of water policy and management. We integrate numerical weather prediction ensembles, distributed hydrological model, and machine learning to generate ensemble streamflow forecasts at medium-range lead times (1–7 days). We demonstrate the application of machine learning as postprocessor for improving the quality of ensemble streamflow forecasts. Our results show that the machine learning postprocessor can improve streamflow forecasts relative to low-complexity forecasts (e.g., climatological and temporal persistence) as well as standalone hydrometeorological modeling and neural network. The relative gain in forecast skill from postprocessor is generally higher at medium-range timescales compared to shorter lead times; high flows compared to low–moderate flows, and the warm season compared to the cool ones. Overall, our results highlight the benefits of machine learning in many aspects for improving both the skill and reliability of streamflow forecasts.

54 ENVIRONMENTAL SCIENCES↗

Comparison of structurally diverse simulation models for prediction of epidemic outcomes caused by a long-distance dispersed pathogen

Long-distance dispersal (LDD) pathogens pose substantial challenges for epidemic control due to their ability to generate new infection foci at great distances. While various modeling approaches have been developed to understand and manage such outbreaks, little work has compared how models of different structures behave under shared conditions. Here, in this study, we compare four structurally distinct epidemiological models — EPIMUL, GEMF, PoPS, and Warwick — each adapted to simulate the spread of wheat stripe rust (WSR), a wind-dispersed LDD pathogen, under identical epidemiological parameters and dispersal kernel. Using data from a controlled field experiment, we evaluate the ability of each model to replicate disease prevalence under nine intervention scenarios that vary in timing and culling area. While the models differ substantially in design — ranging from spatial grid-based to network-based and raster-based frameworks — the shared dispersal kernel allowed for close alignment in their predictions. All models accurately captured general epidemic trends, particularly the strong effect of early intervention on disease suppression. We qualitatively compared their behavioral responses across scenarios and also evaluated an ensemble prediction by averaging across model outputs. Our findings highlight how integrating shared epidemiological components into distinct modeling frameworks can improve consistency and accuracy, while reinforcing the importance of early culling in managing LDD pathogen outbreaks.

Dispersal kernel↗

Improving Enzyme Optimum Temperature Prediction with Resampling Strategies and Ensemble Learning

Accurate prediction of the optimal catalytic temperature ( T opt ) of enzymes is vital in biotechnology, as enzymes with high T opt values are desired for enhanced reaction rates. Recently, a machine learning method (temperature optima for microorganisms and enzymes, TOME) for predicting T opt was developed. TOME was trained on a normally distributed data set with a median T opt of 37 °C and less than 5% of T opt values above 85 °C, limiting the method’s predictive capabilities for thermostable enzymes. Due to the distribution of the training data, the mean squared error on T opt values greater than 85 °C is nearly an order of magnitude higher than the error on values between 30 and 50 °C. Here, we apply ensemble learning and resampling strategies that tackle the data imbalance to significantly decrease the error on high T opt values (>85 °C) by 60% and increase the overall R 2 value from 0.527 to 0.632. The revised method, temperature optima for enzymes with resampling (TOMER), and the resampling strategies applied in this work are freely available to other researchers as Python packages on GitHub.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cybersecurity Anomaly Detection in SCADA-Assisted OT Networks Using Ensemble-Based State Prediction Model

The cybersecurity threats of power system gradually grow due to the increased sophisticated interactions between Information Technology (IT) and Operational Technology (OT) networks. False data injection attack (FDIA) that aims to compromise the Supervisory Control and Data Acquisition (SCADA) measurement and disturb the system operation is one of such cyber threats. Such attacks can potentially lead to significant operational issues at the control centers and substations, and hence, result in severe physical consequences. To avoid catastrophic failure across the power grid resulting from these attacks, it is essential to arm the OT network with real-time vulnerability assessment tools. To this end, this paper outlines various drawbacks of the Purdue architecture model to defend against cyberattacks in the OT network. Furthermore, a novel ensemble-based state prediction model is proposed to detect cybersecurity anomalies in SCADA assisted OT networks. The proposed model uses control center level generation and load forecasts, scheduled, and forced outages, power flow solutions, and the substation level historical data. The hypothesis of the proposed scheme relies on the fact that additional control center and substation data can hardly be accessed and compromised by attackers. One of the vital features of the proposed scheme is an hour-ahead prediction of the operational feasibility of the SCADA measurement range at the control center and substation in real time helps in detecting anomalies in measurements across both substation and the control center.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

High dimensional predictions of suicide risk in 4.2 million US Veterans using ensemble transfer learning

We present an ensemble transfer learning method to predict suicide from Veterans Affairs (VA) electronic medical records (EMR). A diverse set of base models was trained to predict a binary outcome constructed from reported suicide, suicide attempt, and overdose diagnoses with varying choices of study design and prediction methodology. Each model used twenty cross-sectional and 190 longitudinal variables observed in eight time intervals covering 7.5 years prior to the time of prediction. Ensembles of seven base models were created and fine-tuned with ten variables expected to change with study design and outcome definition in order to predict suicide and combined outcome in a prospective cohort. The ensemble models achieved c-statistics of 0.73 on 2-year suicide risk and 0.83 on the combined outcome when predicting on a prospective cohort of ~4.2 M veterans. The ensembles rely on nonlinear base models trained using a matched retrospective nested case-control (Rcc) study cohort and show good calibration across a diversity of subgroups, including risk strata, age, sex, race, and level of healthcare utilization. In addition, a linear Rcc base model provided a rich set of biological predictors, including indicators of suicide, substance use disorder, mental health diagnoses and treatments, hypoxia and vascular damage, and demographics. Similar content being viewed by others

60 APPLIED LIFE SCIENCES↗

Plateau to River Model Predictive Simulations for All Ensemble Realizations to Support Modeling Work in Fiscal Year 2025

The purpose of this environmental calculation file (ECF) is to document predictions of flow and hydraulic head on the Central Plateau of the Hanford Site using the Plateau-to-River (P2R) Model (CP-57037, Model Package Report for the Plateau-to-River Model: Version 9.1). This calculation documents the simulation of the groundwater for the parent model domain of the P2R Model as a basis for use in other applications of the P2R Model. This application is unique from the standpoint that it will simulate all ensemble member models of the P2R Model whereas other applications may only utilize specific ensemble members. These simulations provide results that can be used in the process of selecting an appropriate subset of ensemble members for other applications.

54 ENVIRONMENTAL SCIENCES↗

Data from: Coupled machine learning-ecosystem ensemble models substantially improve predictions of nitrous oxide (N 2 O) fluxes from US croplands

Nitrous oxide (N₂O) is a potent and persistent greenhouse gas, with rising atmospheric concentrations driven in part by inefficient use of synthetic nitrogen (N) fertilizers in agriculture. Predicting soil N₂O emissions is challenging due to high spatial and temporal variability arising from complex soil biogeochemical processes. Process-based ecosystem models and standalone machine learning (ML) approaches without extensive site-specific calibration often miss high emission episodes. Here, we show how an Ensemble Modeling System (EMS) based on outputs from an ensemble of ecosystem models coupled to an ensemble of ML models can improve predictions and understanding of N2O fluxes from US cropland. Trained and validated on approximately 12,000 N2O chamber measurements at 17 U.S. Midwest sites (six crops, 35 management practices), the EMS accurately predicted daily fluxes of N2O at both training (R² = 0.84, RMSE = 16.4 g N ha⁻¹ d⁻¹) and held-out testing sites (R² = 0.84, RMSE = 6.2 g N ha⁻¹ d⁻¹). Analyses identified six dominant N₂O drivers: soil organic carbon (SOC), NH₄⁺, NO₃⁻, water-filled pore space (WFPS), soil temperature, and biomass production. Wet, warm soils produced large N₂O peaks only with sufficient SOC and mineral N; in low-SOC soils, fluxes remained low. Incorporating these drivers into process-based models might significantly improve their predictive capacity. The EMS demonstrates a strong potential to predict N₂O fluxes at unseen sites, enabling more reliable regional inventories, improved gap-filling where measurements are sparse, and enhanced understanding of mechanisms to advance targeted mitigation strategies in food, feed, and bioenergy crops.

agricultural sciences↗