Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Insight into molecular basis and dynamics of full-length CRaf kinase in cellular signaling mechanisms

Raf kinases play key roles in signal transduction in cells for regulating proliferation, differentiation, and survival. Despite decades of research into functions and dynamics of Raf kinases with respect to other cytosolic proteins, understanding Raf kinases is limited by the lack of their full-length structures at the atomic resolution. Here, we present the first model of the full-length CRaf kinase obtained from artificial intelligence/machine learning algorithms with a converging ensemble of structures simulated by large-scale temperature replica exchange simulations. Our model is validated by comparing simulated structures with the latest cryo-EM structure detailing close contacts among three key domains and regions of the CRaf. Our simulations identify potentially new epitopes of intramolecule interactions within the CRaf and reveal a dynamical nature of CRaf kinases, in which the three domains can move back and forth relative to each other for regulatory dynamics. The dynamic conformations are then used in a docking algorithm to shed insight into the paradoxical effect caused by vemurafenib in comparison with a paradox breaker PLX7904. In this study, we propose a model of Raf-heterodimer/KRas-dimer as a signalosome based on the dynamics of the full-length CRaf.

59 BASIC BIOLOGICAL SCIENCES↗

A Bayesian Multi-fidelity Neural Network to Predict Nonlinear Frequency Backbone Curves

The use of structural mechanics models during the design process often leads to the development of models of varying fidelity. Often low-fidelity models are efficient to simulate but lack accuracy, while the high-fidelity counterparts are accurate with less efficiency. Here, this paper presents a multi-fidelity surrogate modeling approach that combines the accuracy of a high-fidelity finite element model with the efficiency of a low-fidelity model to train an even faster surrogate model that parameterizes the design space of interest. The objective of these models is to predict the nonlinear frequency backbone curves of the Tribomechadynamics Research Challenge benchmark structure which exhibits simultaneous nonlinearities from frictional contact and geometric nonlinearity. The surrogate model consists of an ensemble of neural networks that learn the mapping between low and high-fidelity data through nonlinear transformations. Bayesian neural networks are used to assess the surrogate model's uncertainty. Once trained, the multi-fidelity neural network is used to perform sensitivity analysis to assess the influence of the design parameters on the predicted backbone curves. Additionally, Bayesian calibration is performed to update the input parameter distributions to correlate the model parameters to the collection of experimentally measured backbone curves.

42 ENGINEERING↗

Quantifying the Known Unknown: Including Marine Sources of Greenhouse Gases in Climate Modeling

Researchers have recently estimated that Arctic submarine permafrost currently traps 60 billion tons of methane and contains 560 billion tons of organic carbon in seafloor sediments and soil, a giant pool of carbon with potentially large feedbacks on the climate system. Unlike terrestrial permafrost, the submarine permafrost system has remained a “known unknown” because of the difficulty in acquiring samples and measurements. Consequently, this potentially large carbon stock never yet considered in global climate models or policy discussions, represents a real wildcard in our understanding of Earth’s climate. This report summarizes our group’s effort at developing a numerical modeling framework designed to produce a first-of-its-kind estimate of Arctic methane gas releases from the marine sediments to the water column, and potentially to the atmosphere, where positive climate feedback may occur. Newly developed modeling capability supported by the Laboratory Directed Research and Development (LDRD) program at Sandia National Laboratories now gives us the ability to probabilistically map gas distribution and quantity in the seabed by using a hybrid approach of geospatial machine learning, and predictive numerical thermodynamic ensemble modeling. The novelty in this approach is its ability to produce maps of useful data in regions that are only sparsely sampled, a common challenge in the Arctic, and a major obstacle to progress in the past. By applying this model to the circum-Arctic continental shelves and integrating the flux of free gas from in situ methanogenesis and dissociating gas hydrates from the sediment column under climate forcing, we can provide the most reliable estimate of a spatially and temporally varying source term for greenhouse gas flux that can be used by global oceanographic circulation and Earth system models (such as DOE’s E3SM). The result will allow us to finally tackle the wildcard of the submarine permafrost carbon system, and better inform us about the severity of future national security threats that sustained climate change poses.

54 ENVIRONMENTAL SCIENCES↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗

Nonlinear Ensemble Filtering with Diffusion Models: Application to the Surface Quasigeostrophic Dynamics

The intersection between classical data assimilation methods and novel machine learning techniques has attracted significant interest in recent years. Here, we explore another promising solution in which diffusion models are used to formulate a robust nonlinear ensemble filter for sequential data assimilation. Unlike standard machine learning methods, the proposed ensemble score filter (EnSF) is completely training free and can efficiently generate a set of analysis ensemble members. Here, in this study, we apply the EnSF to a surface quasigeostrophic model and compare its performance against the popular local ensemble transform Kalman filter (LETKF), which makes Gaussian assumptions in the analysis step. Numerical tests demonstrate that EnSF maintains stable performance in the absence of localization and for a variety of experimental settings. We find that while LETKF maintains optimal performance in the case of linear observations of the entire state and a perfect model, EnSF shows improvements over LETKF when nonlinear observations are assimilated and the system is subject to unexpected model errors. A spectral decomposition of the analysis results in this nonlinear observation regime shows that the largest improvements over LETKF occur at large scales (small wavenumbers), where LETKF lacks sufficient ensemble spread. Overall, this initial application of EnSF to a geophysical model of intermediate complexity motivates further development of the algorithm for more realistic problems.

Artificial intelligence↗

Accelerated Depth Computation for Surface Boxplots with Deep Learning

Functional depth is a well-known technique used to derive descriptive statistics (e.g., median, quartiles, and outliers) for 1D data. Surface boxplots extend this concept to ensembles of images, helping scientists and users identify representative and outlier images. However, the computational time for surface boxplots increases cubically with the number of ensemble members, making it impractical for integration into visualization tools. In this paper, we propose a deep-learning solution for efficient depth prediction and computation of surface boxplots for time-varying ensemble data. Our deep learning framework accurately predicts member depths in a surface boxplot, achieving average speedups of 6X on a CPU and 15X on a GPU for the 2D Red Sea dataset with 50 ensemble members compared to the traditional depth computation algorithm. Our approach achieves at least a 99% level of rank preservation, with order flipping occurring only at pairs with extremely similar depth values that pose no statistical differences. This local flipping does not significantly impact the overall depth order of the ensemble members.

Han, Mengjiao↗

HydraGNN_GFM_FineTuning4Materials v1.0

This repository enables fine-tuning of the HydraGNN Predictive GFM 2026 — an open-source ensemble of pre-trained graph foundation models for atomistic materials modeling, developed at Oak Ridge National Laboratory. The GFM 2026 is freely available and downloadable via Globus from the OLCF Data Constellation (DOI: 10.13139/OLCF/2562660). Starting from these pre-trained weights, this repository provides a complete transfer learning pipeline for adapting the GFM ensemble to domain-specific molecular and materials property prediction tasks. It includes: 1) Utilities for ensemble fine-tuning with task-specific output heads 2) Example pipelines for eight widely-used materials and molecular datasets 3) Tools for model adaptation and head configuration 4) Data preprocessing utilities for each supported dataset 5) Benchmarking and evaluation scripts

Ungerboeck, Linda↗

Fouling modeling and prediction approach for heat exchangers using deep learning

In this article, we develop a generalized and scalable statistical model for accurate prediction of fouling resistance using commonly measured parameters of industrial heat exchangers. This prediction model is based on deep learning where a scalable algorithmic architecture learns non-linear functional relationships between a set of target and predictor variables from large number of training samples. Here, the efficacy of this modeling approach is demonstrated for predicting fouling in an analytically modeled cross-flow heat exchanger, designed for waste heat recovery from flue-gas using room temperature water. The performance results of the trained models demonstrate that the mean absolute prediction errors are under 10 –4 KW –1 for flue-gas side, water side and overall fouling resistances. The coefficients of determination (R 2 ), which characterize the goodness of fit between the predictions and observed data, are over 99%. Even under varying levels of measurement noise in the inputs, we demonstrate that predictions over an ensemble of multiple neural networks achieves better accuracy and robustness to noise. We find that the proposed deep-learning fouling prediction framework learns to follow heat exchanger flow and heat transfer physics, which we confirm using locally interpretable model agnostic explanations around randomly selected operating points. Overall, we provide a robust algorithmic framework for fouling prediction that can be generalized and scaled to various types of industrial heat exchangers.

42 ENGINEERING↗

Power System Event Classification and Localization Using a Convolutional Neural Network

Detection and timely identification of power system disturbances are essential for situation awareness and reliable electricity grid operation. Because records of actual events in the system are limited, ensemble simulation-based events are needed to provide adequate data for building event-detection models through deep learning; e.g., a convolutional neural network (CNN). An ensemble numerical simulation-based training data set have been generated through dynamic simulations performed on the Polish system with various types of faults in different locations. Such data augmentation is proven to be able to provide adequate data for deep learning. The synchronous generators’ frequency signals are used and encoded into images for developing and evaluating CNN models for classification of fault types and locations. With a time-domain stacked image set as the benchmark, two different time-series encoding approaches, i.e., wavelet decomposition-based frequency-domain stacking and polar coordinate system-based Gramian Angular Field (GAF) stacking, are also adopted to evaluate and compare the CNN model performance and applicability. The various encoding approaches are suitable for different fault types and spatial zonation. With optimized settings of the developed CNN models, the classification and localization accuracies can go beyond 84 and 91%, respectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Domain knowledge-informed, process-mapping AI graph for designing Fe-based alloys

<span style="font-family: Calibri, sans-serif; font-size: 12pt;">Continuous improvement in efficiency of a power plant relies on designing materials for use at increasingly higher temperature and/or pressure, for 100,000s hours of operation. Due to complexity, non-linearity and high-dimensionality of the problem, traditional Machine Learning (ML) approaches require unreasonably large datasets for the data-driven model development. Science-based material and process engineering complements hard data with, sometimes soft and intuitive, empirical domain knowledge. Artificial Intelligence (AI) was used in this study to incorporate such knowledge into computational graph architecture (process-mimicking artificial neuron design, causal layer and graph structures, ensemble modeling of latent states) and learning procedures (variable transformation, fuzzy physics pre-training and freezing of deep layers, virtual microstructure representation, and adversarial multi-objective optimization). The first alloys design pathways suggested by the AI tool (pyroMind) passed a preliminary engineering review on soundness and transparency.</span>

Romanov, Vyacheslav↗

Autodifferentiable Ensemble Kalman Filters

Data assimilation is concerned with sequentially estimating a temporally evolving state. This task, which arises in a wide range of scientific and engineering applications, is particularly challenging when the state is high-dimensional and the state-space dynamics are unknown. This paper introduces a machine learning framework for learning dynamical systems in data assimilation. Here, our auto-differentiable ensemble Kalman filters (AD-EnKFs) blend ensemble Kalman filters for state recovery with machine learning tools for learning the dynamics. In doing so, AD-EnKFs leverage the ability of ensemble Kalman filters to scale to high-dimensional states and the power of automatic differentiation to train high-dimensional surrogate models for the dynamics. Numerical results using the Lorenz-96 model show that AD-EnKFs outperform existing methods that use expectation-maximization or particle filters to merge data assimilation and machine learning. In addition, AD-EnKFs are easy to implement and require minimal tuning.

autodifferentiation↗

Assessing decadal variability of subseasonal forecasts of opportunity using explainable AI

Abstract Identifying predictable states of the climate system allows for enhanced prediction skill on the generally low-skill subseasonal timescale via forecasts with higher confidence and accuracy, known as forecasts of opportunity. This study takes a neural network approach to explore decadal variability of subseasonal predictability, particularly during forecasts of opportunity. Specifically, this work quantifies subseasonal prediction skill provided by the tropics within the Community Earth System Model Version 2 (CESM2) Large Ensemble and assesses how this skill evolves on decadal timescales. Utilizing the networks’ confidence and explainable artificial intelligence, physically meaningful sources of predictability associated with periods of enhanced skill are identified. Using these networks, we find that tropically-driven subseasonal predictability varies on decadal timescales during forecasts of opportunity. Further, we investigate the drivers of the low frequency modulation of the tropical-extratropical teleconnection and discuss the implications. Analysis is extended to ECMWF Reanalysis v5 data, revealing that the relationships learned within the CESM2-Large Ensemble holds in modern reanalysis data. These results indicate that the neural networks are capable of identifying predictable decadal states of the climate system within CESM2 that are useful for making confident, accurate subseasonal precipitation predictions in the real world.

54 ENVIRONMENTAL SCIENCES↗

Structure–activity relationship-based chemical classification of highly imbalanced Tox21 datasets

Abstract The specificity of toxicant-target biomolecule interactions lends to the very imbalanced nature of many toxicity datasets, causing poor performance in Structure–Activity Relationship (SAR)-based chemical classification. Undersampling and oversampling are representative techniques for handling such an imbalance challenge. However, removing inactive chemical compound instances from the majority class using an undersampling technique can result in information loss, whereas increasing active toxicant instances in the minority class by interpolation tends to introduce artificial minority instances that often cross into the majority class space, giving rise to class overlapping and a higher false prediction rate. In this study, in order to improve the prediction accuracy of imbalanced learning, we employed SMOTEENN, a combination of Synthetic Minority Over-sampling Technique (SMOTE) and Edited Nearest Neighbor (ENN) algorithms, to oversample the minority class by creating synthetic samples, followed by cleaning the mislabeled instances. We chose the highly imbalanced Tox21 dataset, which consisted of 12 in vitro bioassays for > 10,000 chemicals that were distributed unevenly between binary classes. With Random Forest (RF) as the base classifier and bagging as the ensemble strategy, we applied four hybrid learning methods, i.e., RF without imbalance handling (RF), RF with Random Undersampling (RUS), RF with SMOTE (SMO), and RF with SMOTEENN (SMN). The performance of the four learning methods was compared using nine evaluation metrics, among which F 1 score, Matthews correlation coefficient and Brier score provided a more consistent assessment of the overall performance across the 12 datasets. The Friedman’s aligned ranks test and the subsequent Bergmann-Hommel post hoc test showed that SMN significantly outperformed the other three methods. We also found that a strong negative correlation existed between the prediction accuracy and the imbalance ratio (IR), which is defined as the number of inactive compounds divided by the number of active compounds. SMN became less effective when IR exceeded a certain threshold (e.g., > 28). The ability to separate the few active compounds from the vast amounts of inactive ones is of great importance in computational toxicology. This work demonstrates that the performance of SAR-based, imbalanced chemical toxicity classification can be significantly improved through the use of data rebalancing.

Idakwo, Gabriel↗

Machine learning visualization tool for exploring parameterized hydrodynamics

We are interested in the computational study of shock hydrodynamics, i.e. problems involving compressible solids, liquids, and gases that undergo large deformation. These problems are dynamic and nonlinear and can exhibit complex instabilities. Due to advances in high performance computing it is possible to parameterize a hydrodynamic problem and perform a computational study yielding $\mathscr{O} (TB)$ of simulation state data. We present an interactive machine learning tool that can be used to compress, browse, and interpolate these large simulation datasets. This tool allows computational scientists and researchers to quickly visualize 'what-if' situations, perform sensitivity analyses, and optimize complex hydrodynamic experiments.

97 MATHEMATICS AND COMPUTING↗

Inferring safety critical events from vehicle kinematics in naturalistic driving environment: Application of deep learning Algorithms

Advances in sensing technology has enabled the collection of countless terabytes of second-by-second kinematics data. Such data provides opportunities for real-time monitoring of driving behavior and identification of safety critical events (SCEs) including crashes and near crashes. The concept of volatility is relevant in this context, which identifies instability and erratic variations in driving behavior prior to involvement in SCEs. This study utilized vehicle kinematics from a large-scale naturalistic driving data to develop a deep learning approach based on 1D convolutional neural networks (CNN) for inferring SCEs. The data are unique in the sense that such accurate pre-crash data at high fidelity are not available in traditional crash repositories. This study contributes to the literature by providing a first attempt at predicting responses to SCEs by developing deep learning-based CNN architectures using novel driving volatility based kinematic thresholds for a sample of 9553 events. The key contribution lies in developing a volatility-based CNN input layout that is acceptable to CNN schemes and represents the motion kinematics such as speed, acceleration and volatility measures. Several 1D-CNN architectures were developed using layers, numbers of convolutions, layer patterns, and kernels. Shallow and deep architectures were tested, revealing higher accuracy of shallow architectures in detecting SCEs. The optimal number of epochs were identified using an early stopping method while the CNN performance was improved by increasing the number of epochs. The ensemble CNN had the highest predictive accuracy of 95.6% for detection of crashes and near crashes, which was 2.5% higher than the optimal CNN using 20% hold out test data. The ensemble CNN also outperformed classical machine learning models and model performance reported in past studies on detection of SCEs. Finally, these results have implications for identification of safety hotspots and providing real-time alerts and warnings in connected and highly automated vehicle environment including society of automotive engineers levels 3–5.

42 ENGINEERING↗

Progress in Normalizing Flows for 4d Gauge Theories

Normalizing flows have arisen as a tool to accelerate Monte Carlo sampling for lattice field theories. This work reviews recent progress in applying normalizing flows to 4-dimensional nonabelian gauge theories, focusing on two advancements: an architectural improvement referred to as learned active loops, and the application of correlated ensemble methods to QCD with N f = 2 dynamical fermions.

Abbott, Ryan [Massachusetts Institute of Technolog↗

Predicting Biomass Yields of Advanced Switchgrass Cultivars for Bioenergy and Ecosystem Services Using Machine Learning

The production of advanced perennial bioenergy crops within marginal areas of the agricultural landscape is gaining interest due to its potential to sustainably produce feedstocks for biofuels and bioproducts while also improving the sustainability and resilience of commodity crop production. However, predicting the biomass yields of this production system is challenging because marginal areas are often relatively small and spread around agricultural fields and are typically associated with various abiotic conditions that limit crop production. Machine learning (ML) offers a viable solution as a biomass yield prediction tool because it is suited to predicting relationships with complex functional associations. The objectives of this study were to (1) evaluate the accuracy of commonly applied ML algorithms in agricultural applications for predicting the biomass yields of advanced switchgrass cultivars for bioenergy and ecosystem services and (2) determine the most important biomass yield predictors. Datasets on biomass yield, weather, land marginality, soil properties, and agronomic management were generated from three field study sites in two U.S. Midwest states (Illinois and Iowa) over three growing seasons. The ML algorithms evaluated in the study included random forests (RFs), gradient boosting machines (GBMs), artificial neural networks (ANNs), K-neighbors regressor (KNR), AdaBoost regressor (ABR), and partial least squares regression (PLSR). Coefficient of determination (R 2 ) and mean absolute error (MAE) were used to evaluate the predictive accuracy of the tested algorithms. Results showed that the ensemble methods, RF (R 2 = 0.86, MAE = 0.62 Mg/ha), GBM (R 2 = 0.88, MAE = 0.57 Mg/ha), and GBM (R 2 = 0.78, MAE = 0.66 Mg/ha), were the most accurate in predicting biomass yields of the Independence, Liberty, and Shawnee switchgrass cultivars, respectively. This is in agreement with similar studies that apply ML to multi-feature problems where traditional statistical methods are less applicable and datasets used were considered to be relatively small for ANNs. Consistent with previous studies on switchgrass, the most important predictors of biomass yield included average annual temperature, average growing season temperature, sum of the growing season precipitation, field slope, and elevation. This study helps pave the way for applying ML as a management tool for alternative bioenergy landscapes where understanding agronomic and environmental performance of a multifunctional cropping system seasonally and interannually at the sub-field scale is critical.

09 BIOMASS FUELS↗