Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Peatland fires in Alaska will double by the end of the century

During recent summers, warm and dry conditions have increased the occurrence of wildfires and potentially peat-fires across Alaska. Limitations in resolving the fine-scale distribution of peatlands and climate observations have constrained our ability to accurately predict peat-fire dynamics. Using a new high-resolution peatland map of Alaska, we evaluated the climate and environmental controls of past and future peat-fire activity. Ensemble machine learning models identified reduced soil moisture, higher temperatures, and evapotranspiration as key predictors of annual total burned peatland area (tenfold CV R 2 = 0.62, RMSE = 221.1 km 2 ). By the end of the twenty-first century, models forced with climate datasets from representative concentration pathways (RCPs) 4.5, 6.0, and 8.5 emission scenarios project a statewide doubling of burned peatlands (increasing 61–121%), with regional increases ranging from 25–165% in polar, 61–95% in boreal, and 102–106% in maritime ecoregions. These projections indicate that wildfires will progressively encroach further into organic-rich moist and wet peaty soils, potentially amplifying soil carbon release across Alaska.

climate-change ecology↗

Identifications of RR Lyrae Stars and Quasars from the Simulated Data of Mephisto-W Survey

We have investigated the feasibilities and accuracies of the identifications of RR Lyrae stars and quasars from the simulated data of the Multi-channel Photometric Survey Telescope (Mephisto) W Survey. Based on the variable sources light curve libraries from the Sloan Digital Sky Survey (SDSS) Stripe 82 data and the observation history simulation from the Mephisto-W Survey Scheduler, we have simulated the uvgriz multi-band light curves of RR Lyrae stars, quasars and other variable sources for the first-year observation of Mephisto W Survey. We have applied the ensemble machine learning algorithm Random Forest Classifier (RFC) to identify RR Lyrae stars and quasars, respectively. We build training and test samples and extract ~150 features from the simulated light curves and train two RFCs respectively for the RR Lyrae star and quasar classification. We find that, our RFCs are able to select the RR Lyrae stars and quasars with remarkably high precision and completeness, with purity = 95.4% and completeness = 96.9% for the RR Lyrae RFC and purity = 91.4% and completeness = 90.2% for the quasar RFC. In conclusion, we have also derived relative importances of the extracted features utilized to classify RR Lyrae stars and quasars.

(galaxies:) quasars: general↗

Navigating Uncertainty: Challenges in Visualizing Ensemble Data and Surrogate Models for Decision Systems

Uncertainty visualization plays a critical role in transforming ensemble simulation data into actionable insights by effectively communicating various dimensions of uncertainty within a system. The emergence of artificial intelligence-driven surrogate models trained on multirun ensemble data offers a transformative opportunity to replace computationally intensive simulations with fast estimates, enabling users to explore data spaces with unprecedented depth and interactivity. However, integrating ensemble data and surrogate models into decision-making workflows and tools introduces novel challenges for uncertainty visualization. These include reconciling and clearly communicating the unique uncertainties associated with ensembles and their surrogate model estimates, and leveraging these approximations to inform actionable decisions. This work explores these challenges in the context of high-dimensional data visualization, bridging discrete datasets with their continuous representations and addressing the complexities of systems that support iterative navigation between input and output spaces. We evaluate the role of uncertainty visualization in fostering intuitive, actionable interactions and identify critical hurdles in advancing this frontier of computational simulation.

97 MATHEMATICS AND COMPUTING↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Continental United States may lose 1.8 petagrams of soil organic carbon under climate change by 2100

Abstract Aims High‐resolution information on soils’ vulnerability to climate‐induced soil organic carbon (SOC) loss can enable environmental scientists, land managers, and policy makers to develop targeted mitigation strategies. This study aims to estimate baseline and decadal changes in continental US surface SOC stocks under future emission scenarios. Location Continental United States. Time period 2014–2100. Methods We used recent SOC field observations ( n = 6,213 sites), environmental factors ( n = 32), and an ensemble machine learning (ML) approach to estimate baseline SOC stocks in surface soils across the continental United States at 100‐m spatial resolution, and decadal changes under the projected climate scenarios of Coupled Model Intercomparison Project Phase Six (CMIP6) earth system models (ESMs). Results Baseline SOC projections from ML approaches captured more than 50% of variability in SOC observations, whereas ESMs represented only 6–16% of observed SOC variability. ML estimates showed a mean total loss of 1.8 Pg C from US surface soils under the high‐emission scenario by 2100, whereas ESMs showed no significant change in SOC stocks with wide variation among ESMs. Both ML and ESM predictions agree on the direction of SOC change (net emissions or sequestration) across 46–51% of continental US land area. These differences are attributable to the high‐resolution site‐specific data used in the ML models compared to the relatively coarse grid represented in CMIP6 ESMs. Main conclusions Our high‐resolution estimates of baseline SOC stocks, identification of key environmental controllers, and projection of SOC changes from US land cover types under future climate scenarios suggest the need for high‐resolution simulations of SOC in ESMs to represent the heterogeneity of SOC. We found that the SOC change is sensitive to key soil related factors (e.g. soil drainage and soil order) that have not been historically considered as input parameters in ESMs, because currently more than 95% variability in the SOC of CMIP6 ESMs is controlled by net primary productivity, temperature, and precipitation. Using additional environmental factors to estimate the baseline SOC stocks and predict the future trajectory of SOC change can provide more accurate results.

54 ENVIRONMENTAL SCIENCES↗

Identifying precursors of daily to seasonal hydrological extremes over the USA using deep learning techniques and climate model ensembles

Focal Area(s): We focus on two areas of crosscutting interest for DOE: 1) predictability of extreme precipitation and drought in the USA and 2) the integration of climate models with new AI tools, such as convolutional neural networks (CNN) and methods to understand their output (e.g. layer-wise relevance propagation; LRP). This project fits into focus area 3 of this call for white paper using AI to gain insight from complex data, including explainable AI tools. Science Challenge: Predicting hydrological extremes is important due to their impacts on people, agriculture and infrastructure. This prediction is difficult due to the infrequent occurrence of extremes and their complexity. However, extreme events can be related to more predictable conditions in the ocean, such as El Nino, long-term soil moisture or large scale modes of climate variability, such as the North Atlantic Oscillation (NAO).

54 ENVIRONMENTAL SCIENCES↗

Constraining microphysical processes of warm rain formulation using advanced spectral separations, an ensemble retrieval framework and machine learning techniques

Drizzle, a common feature of marine boundary layer clouds formed through collision coalescence, plays a key role in cloud microphysics and evolution. Yet, simultaneously retrieving cloud and drizzle properties from remote-sensing observations remains challenging because drizzle droplets often dominate radar signals, masking cloud contributions. The goal of the proposed research is to provide constraints for the process of autoconversion and accretion using ARM cloud measurements. Specifically, we provide concurrent retrievals of cloud and drizzle that allows users to derive corresponding autoconversion and accretion rates.

54 ENVIRONMENTAL SCIENCES↗

Robust Explanations using Diverse Adversarially Trained Ensembles, Multi-Modal Contrastive Learning, and Attribution-based Confidence Metrics

The primary objective of this project is to strengthen the trustworthiness of AI systems by designing algorithms that make their internal decision-making processes more understandable to human users. This involves creating clear, interpretable explanations for AI decisions and developing metrics to assess these explanations' validity and reliability. Significant progress has been achieved through (i) developing symbolic explanations, (ii) generating meaningful interpretive insights, (iii) establishing accuracy and confidence metrics, and (iv) devising methods to evaluate the knowledge boundaries of AI models. To date, the research findings have been shared in peer-reviewed publications, with accompanying scientific and technical information (STI) detailed below.

97 MATHEMATICS AND COMPUTING↗

Navigating the Noise: Bringing Clarity to ML Parameterization Design With O $\boldsymbol{\mathcal{O}}$(100) Ensembles

Abstract Machine‐learning (ML) parameterizations of subgrid processes (here of turbulence, convection, and radiation) may one day replace conventional parameterizations by emulating high‐resolution physics without the cost of explicit simulation. However, uncertainty about the relationship between offline and online performance (i.e., when integrated with a large‐scale general circulation model) hinders their development. Much of this uncertainty stems from limited sampling of the noisy, emergent effects of upstream ML design decisions on downstream online hybrid simulation. Our work rectifies the sampling issue via the construction of a semi‐automated, end‐to‐end pipeline for size ensembles of hybrid simulations, revealing important nuances in how systematic reductions in offline error manifest in changes to online error and online stability. For example, removing dropout and switching from a Mean Squared Error to a Mean Absolute Error loss both reduce offline error, but they have opposite effects on online error and online stability. Other design decisions, like incorporating memory, converting moisture input from specific humidity to relative humidity, using batch normalization, and training on multiple climates do not come with any such compromises. Finally, we show that ensemble sizes of may be necessary to reliably detect causally relevant differences online. By enabling rapid online experimentation at scale, we can empirically settle debates regarding subgrid ML parameterization design that would have otherwise remained unresolved in the noise.

Lin, Jerry [Department of Earth System Sciences Un↗

Integrating an Ensemble Reward System into an Off-Policy Reinforcement Learning Algorithm for the Economic Dispatch of Small Modular Reactor-Based Energy Systems

Nuclear Integrated Energy Systems (NIES) have emerged as a comprehensive solution for navigating the changing energy landscape. They combine nuclear power plants with renewable energy sources, storage systems, and smart grid technologies to optimize energy production, distribution, and consumption across sectors, improving efficiency, reliability, and sustainability while addressing challenges associated with variability. The integration of Small Modular Reactors (SMRs) in NIES offers significant benefits over traditional nuclear facilities, although transferring involves overcoming legal and operational barriers, particularly in economic dispatch. This study proposes a novel off-policy Reinforcement Learning (RL) approach with an ensemble reward system to optimize economic dispatch for nuclear-powered generation companies equipped with an SMR, demonstrating superior accuracy and efficiency when compared to conventional methods and emphasizing RL’s potential to improve NIES profitability and sustainability. Finally, the research attempts to demonstrate the viability of implementing the proposed integrated RL approach in spot energy markets to maximize profits for nuclear-driven generation companies, establishing NIES’ profitability over competitors that rely on fossil fuel-based generation units to meet baseload requirements.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Machine-learning accelerated geometry optimization in molecular simulation

Geometry optimization is an important part of both computational materials and surface science because it is the path to finding ground state atomic structures and reaction pathways. These properties are used in the estimation of thermodynamic and kinetic properties of molecular and crystal structures. This process is slow at the quantum level of theory because it involves an iterative calculation of forces using quantum chemical codes such as density functional theory (DFT), which are computationally expensive and which limit the speed of the optimization algorithms. It would be highly advantageous to accelerate this process because then one could do either the same amount of work in less time or more work in the same time. Here, we provide a neural network (NN) ensemble based active learning method to accelerate the local geometry optimization for multiple configurations simultaneously. We illustrate the acceleration on several case studies including bare metal surfaces, surfaces with adsorbates, and nudged elastic band for two reactions. In all cases, the accelerated method requires fewer DFT calculations than the standard method. In addition, we provide an Atomic Simulation Environment (ASE)-optimizer Python package to make the usage of the NN ensemble active learning for geometry optimization easier.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications

Abstract Robust quantification of predictive uncertainty is a critical addition needed for machine learning applied to weather and climate problems to improve the understanding of what is driving prediction sensitivity. Ensembles of machine learning models provide predictive uncertainty estimates in a conceptually simple way but require multiple models for training and prediction, increasing computational cost and latency. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but does not account for epistemic uncertainty. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainties with one model. This study compares the uncertainty derived from evidential neural networks to that obtained from ensembles. Through applications of the classification of winter precipitation type and regression of surface-layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. To encourage broader adoption of evidential deep learning, we have developed a new Python package, Machine Integration and Learning for Earth Systems (MILES) group Generalized Uncertainty for Earth System Science (GUESS) (MILES-GUESS) ( https://github.com/ai2es/miles-guess ), that enables users to train and evaluate both evidential and ensemble deep learning. Significance Statement This study demonstrates a new technique, evidential deep learning, for robust and computationally efficient uncertainty quantification in modeling the Earth system. The method integrates probabilistic principles into deep neural networks, enabling the estimation of both aleatoric uncertainty from noisy data and epistemic uncertainty from model limitations using a single model. Our analyses reveal how decomposing these uncertainties provides valuable insights into reliability, accuracy, and model shortcomings. We show that the approach can rival standard methods in classification and regression tasks within atmospheric science while offering practical advantages such as computational efficiency. With further advances, evidential networks have the potential to enhance risk assessment and decision-making across meteorology by improving uncertainty quantification, a longstanding challenge. This work establishes a strong foundation and motivation for the broader adoption of evidential learning, where properly quantifying uncertainties is critical yet lacking.

Schreck, John S.↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

Using ensembles and distillation to optimize the deployment of deep learning models for the classification of electronic cancer pathology reports

One of the goals of the Surveillance, Epidemiology, and End Results (SEER) program is to estimate incidence, prevalence, and mortality of all cancers. To that end, cancer registries across the country maintain a massive database of cancer pathology reports which contain rich information to understand cancer trends. However, these reports are stored in the form of unstructured text, and human annotators are required to read and extract relevant information. In this article, we show that existing deep learning models for automating information extraction from cancer pathology reports can be significantly improved by using ensemble model distillation. We found that by training multiple predictive models and transferring their knowledge to a single, low-resource model, we can reduce the number of highly confident wrong predictions. Our results show that our implemented methods could save 1000s of manual annotation hours.

60 APPLIED LIFE SCIENCES↗

Accelerated Probabilistic Marching Cubes by Deep Learning for Time-Varying Scalar Ensembles

Visualizing the uncertainty of ensemble simulations is challenging due to the large size and multivariate and temporal features of en-semble data sets. One popular approach to studying the uncertainty of ensembles is analyzing the positional uncertainty of the level sets. Probabilistic marching cubes is a technique that performs Monte Carlo sampling of multivariate Gaussian noise distributions for positional uncertainty visualization of level sets. However, the technique suffers from high computational time, making interactive visualization and analysis impossible to achieve. This paper introduces a deep-learning-based approach to learning the level-set uncertainty for two-dimensional ensemble data with a multivariate Gaussian noise assumption. We train the model using the first few time steps from time-varying ensemble data in our workflow. We demonstrate that our trained model accurately infers uncertainty in level sets for new time steps and is up to 170X faster than that of the original probabilistic model with serial computation and 10X faster than that of the original parallel computation.

Han, Mengjiao↗