Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING↗

Selecting representative geological realizations to model subsurface CO 2 storage under uncertainty

Carbon capture and storage (CCS) is one of the quickest and most effective solutions for reducing carbon emissions. The majority of subsurface storage occurs in saline aquifers, for which geological information is lacking which in turn results in geological uncertainty. To evaluate uncertainty in CO 2 injection projections, the use of multiple geological realizations (GRs) has been practiced very commonly. In this approach, hundreds or thousands of high-resolution GRs is used that quickly becomes computationally expensive. This issue can be addressed with representative geological realizations (RGRs) that preserve the uncertainty domain of the ensemble GRs. Here, in this study, we propose the use of unsupervised machine learning (UML) frameworks, including dissimilarity measurement, dimensionality reduction, clustering and sampling algorithms ta select a predetermined number of RGRs. We compare the simulation outputs of the RGR sets and the ensemble using the Kolmogorov–Smirnov (KS) test to select the best UML. The UML frameworks and their associated selection processes are evaluated using a saline aquifer with a single CO 2 injection well and 200 GRs with varying uncertain petrophysical characteristics. The best UML framework is selected to use only 5% of the GRs while maintaining the uncertainty domain of the ensemble GRs. In addition, the best UML framework is tested using a saline aquifer with three CO 2 injection wells and varied GRs. The results show that our proposed UML framework can be used to choose RGRs, capturing the whole uncertainty domain. Our approach leads to a significant reduction in the computational cost associated with scenario testing, decision-making, and development planning for CO 2 storage sites under geological uncertainty.

58 GEOSCIENCES↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

Cross-Layered Distributed Data-Driven Framework for Enhanced Smart Grid Cyber-Physical Security

Smart Grid (SG) research and development has drawn much attention from academia, industry and government due to the great impact it will have on society, economics and the environment. Securing the SG is a considerably significant challenge due the increased dependency on communication networks to assist in physical process control, exposing them to various cyber-threats. In addition to attacks that change measurement values using False Data Injection (FDI) techniques, attacks on the communication network may disrupt the power system's real-time operation by intercepting messages, or by flooding the communication channels with unnecessary data. Addressing these attacks requires a cross-layer approach. In this paper a cross-layered strategy is presented, called Cross-Layer Ensemble CorrDet with Adaptive Statistics(CECD-AS), which integrates the detection of faulty SG measurement data as well as inconsistent network inter-arrival times and transmission delays for more reliable and accurate anomaly detection and attack interpretation. Numerical results show that CECD-AS can detect multiple False Data Injections, Denial of Service (DoS) and Man In The Middle (MITM) attacks with a high F1-score compared to current approaches that only use SG measurement data for detection such as the traditional physics-based State Estimation, Ensemble CorrDet with Adaptive Statistics strategy and other machine learning classification-based detection schemes.

cyber-physical security↗

Multiscale Machine-Learned Modeling Infrastructure

The Multiscale Machine-Learned Modeling Infrastructure (MuMMI) is a multiscale workflow management infrastructure that can concurrently orchestrate thousands of molecular dynamics (MD) simulations operating at different time and/or length scales, spanning nanoseconds to seconds and nanometers to micrometers. MuMMI uses machine learning (backed by biology experiments) to guide a massive ensemble of MD simulations that capture biologically relevant time and length scales with unprecedented resolution. MuMMI supports multiple MD codes such as GROMACS and ddcMD and can be fully deployed using the HPC package manager Spack. MuMMI has been used in many publications to run hundreds of thousands simulations, leading to significant biology breakthroughs.

Di Natale, Francesco [Lawrence Livermore National ↗

Value of Information and Data Management Platform Promoting Modern Data Analysis Utilization. STTR Phase I, Final Technical Report

The immense growth in weather and river sensor data availability along with increases in computational power and algorithmic sophistication have enabled the widespread application of physical, machine learning, and artificial intelligence algorithms to river forecasting. The forecasts produced by widely varying model approaches are inherently limited in different ways, such as by the number and quality of training observations, producing biased predictions that achieve optimal skill only over specific ranges of weather and river conditions. Thus, despite the growing availability and new sources of forecasts, hydropower owners cannot fully leverage the information because they lack an effective mechanism to determine the “best” forecast for given conditions or objectives. Tools are needed to consolidate forecast information, demonstrate skill, and extract optimal forecast guidance from the range of possibilities. Such tools will increase uptake and beneficial use of the growing wealth of forecast information, and lead to improved operating decisions, optimal power generation, and enhanced system value. This research proposes to improve the value of forecasts by developing a broadly accessible verification framework and applying machine learning techniques to identify the ‘best’ forecast from an ensemble of sources given the current weather, river, and operating conditions.

13 HYDRO ENERGY↗

Modeling Weather Impact on Airport Arrival Miles-in-Trail Restrictions

When the demand for either a region of airspace or an airport approaches or exceeds the available capacity, miles-in-trail (MIT) restrictions are the most frequently issued traffic management initiatives (TMIs) that are used to mitigate these imbalances. Miles-intrail operations require aircraft in a traffic stream to meet a specific inter-aircraft separation in exchange for maintaining a safe and orderly flow within the stream. This stream of aircraft can be departing an airport, over a common fix, through a sector, on a specific route or arriving at an airport. This study begins by providing a high-level overview of the distribution and causes of arrival MIT restrictions for the top ten airports in the United States. This is followed by an in-depth analysis of the frequency, duration and cause of MIT restrictions impacting the Hartsfield-Jackson Atlanta International Airport (ATL) from 2009 through 2011. Then, machine-learning methods for predicting (1) situations in which MIT restrictions for ATL arrivals are implemented under low demand scenarios, and (2) days in which a large number of MIT restrictions are required to properly manage and control ATL arrivals are presented. More specifically, these predictions were accomplished by using an ensemble of decision trees with Bootstrap aggregation (BDT) and supervised machine learning was used to train the BDT binary classification models. The models were subsequently validated using data cross validation methods. When predicting the occurrence of arrival MIT restrictions under low demand situations, the model was able to achieve over all accuracy rates ranging from 84% to 90%, with false alarm ratios ranging from 10% to 15%. In the second set of studies designed to predict days on which a high number of MIT restrictions were required, overall accuracy rates of 80% were achieved with false alarm ratios of 20%. Overall, the predictions proposed by the model give better MIT usage information than what has been currently provided under current day operations. Traffic flow managers can use these predictions to identify potential MIT restrictions to eliminate (e.g., those occurring during low arrival demand periods), and to determine the days in which a significant number of restrictions may be required

Operation↗

GeoThermalCloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale that has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify highvalue data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies.

58 GEOSCIENCES↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,↗

Unsupervised machine learning for unbiased chemical classification in X-ray absorption spectroscopy and X-ray emission spectroscopy

Here we report a comprehensive computational study of unsupervised machine learning for extraction of chemically relevant information in X-ray absorption near edge structure (XANES) and in valence-to-core X-ray emission spectra (VtC-XES) for classification of a broad ensemble of sulphorganic molecules. By progressively decreasing the constraining assumptions of the unsupervised machine learning algorithm, moving from principal component analysis (PCA) to a variational autoencoder (VAE) to t-distributed stochastic neighbour embedding (t-SNE), we find improved sensitivity to steadily more refined chemical information. Surprisingly, when embedding the ensemble of spectra in merely two dimensions, t-SNE distinguishes not just oxidation state and general sulphur bonding environment but also the aromaticity of the bonding radical group with 87% accuracy as well as identifying even finer details in electronic structure within aromatic or aliphatic sub-classes. We find that the chemical information in XANES and VtC-XES is very similar in character and content, although they unexpectedly have different sensitivity within a given molecular class. We also discuss likely benefits from further effort with unsupervised machine learning and from the interplay between supervised and unsupervised machine learning for X-ray spectroscopies. Our overall results, i.e., the ability to reliably classify without user bias and to discover unexpected chemical signatures for XANES and VtC-XES, likely generalize to other systems as well as to other one-dimensional chemical spectroscopies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A baseline for ensemble-based, time-resolved inflow reconstruction for a single turbine using large-eddy simulations and latent diffusion models

We are interested in reconstructing winds flowing through a turbine on a second-by-second basis over a 10 min window. Previously, we developed a machine learning algorithm that takes in a snapshot of wind speed measurements and generates ensembles of three-dimensional wind field estimates. Here, we use these estimates as initial conditions in large-eddy simulations and reconstruct atmospheric and turbine response dynamic quantities in a synthetic field campaign. In doing so, we establish a baseline for model validation that future time-aware data assimilation techniques will be compared to. In turbine-free case studies, ground truth wind speeds consistently fall within our estimated wind speed distribution for the first 100 s after the simulation start. In simulations with turbines, the wind estimates show a small bias of 0.10 m s 1 and good correlation of 0.80 during the first 100 s. During this window, our estimates of the Blade 1 bending moment and generator power typically span the ground truth, with the estimate of the former performing better overall. In summary, this approach shows promise as a stand-alone technique for reconstructing real-world inflow and turbine dynamics in 1-2 min windows and as a foundation for future time-aware data assimilation techniques.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

DeepHyper: A Python Package for Massively Parallel Hyperparameter Optimization in Machine Learning

Machine learning models are increasingly applied across scientific disciplines, yet their effectiveness often hinges on heuristic decisions—such as data transformations, training strategies, and model architectures—that are not learned by the models themselves. Automating the selection of these heuristics and analyzing their sensitivity is crucial for building robust and efficient learning workflows. DeepHyper addresses this challenge by democratizing hyperparameter optimization, providing accessible tools to streamline and enhance machine learning workflows from a laptop to the largest supercomputer in the world. Building on top of hyperparameter optimization, it unlocks new capabilities around ensembles of models for improved accuracy and uncertainty quantification. All of these organized around efficient parallel computing.

ensemble↗

Ensemble Data Mining Methods

Ensemble Data Mining Methods, also known as Committee Methods or Model Combiners, are machine learning methods that leverage the power of multiple models to achieve better prediction accuracy than any of the individual models could on their own. The basic goal when designing an ensemble is the same as when establishing a committee of people: each member of the committee should be as competent as possible, but the members should be complementary to one another. If the members are not complementary, Le., if they always agree, then the committee is unnecessary---any one member is sufficient. If the members are complementary, then when one or a few members make an error, the probability is high that the remaining members can correct this error. Research in ensemble methods has largely revolved around designing ensembles consisting of competent yet complementary models.

Oza, Nikunj C.↗

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

Galaxy Zoo DECaLS: Detailed visual morphology measurements from volunteers and deep learning for 314,000 galaxies

We present Galaxy Zoo DECaLS: detailed visual morphological classifications for Dark Energy Camera Legacy Survey images of galaxies within the SDSS DR8 footprint. Deeper DECaLS images (r = 23.6 versus r = 22.2 from SDSS) reveal spiral arms, weak bars, and tidal features not previously visible in SDSS imaging. To best exploit the greater depth of DECaLS images, volunteers select from a new set of answers designed to improve our sensitivity to mergers and bars. Galaxy Zoo volunteers provide 7.5 million individual classifications over 314 000 galaxies. 140 000 galaxies receive at least 30 classifications, sufficient to accurately measure detailed morphology like bars, and the remainder receive approximately 5. All classifications are used to train an ensemble of Bayesian convolutional neural networks (a state-of-the-art deep learning method) to predict posteriors for the detailed morphology of all 314 000 galaxies. We use active learning to focus our volunteer effort on the galaxies which, if labelled, would be most informative for training our ensemble. When measured against confident volunteer classifications, the trained networks are approximately 99 per cent accurate on every question. Morphology is a fundamental feature of every galaxy; our human and machine classifications are an accurate and detailed resource for understanding how galaxies evolve.

79 ASTRONOMY AND ASTROPHYSICS↗

Nanoscale wetting controls reactive Pd ensembles in synthesis of dilute PdAu alloy catalysts

The performance of bimetallic dilute alloy catalysts is largely determined by the size of minority metal ensembles on the nanoparticle surface. By analyzing the synthesis of catalysts comprising Pd 8 Au 92 nanoparticles supported on silica using surface-sensitive techniques, we report that whether Pd overgrowth occurs before or after Au nanoparticle deposition onto the support controls the surface Pd ensemble size and abundance. These differences in Pd ensembles influence catalytic reactivity in H 2 –D 2 isotope exchange and benzaldehyde hydrogenation, which, in correlation with theoretical calculations, is used to elucidate the active site(s) in each reaction. To clarify how the synthetic sequence controls the formation of Pd ensembles, we combine numerical wetting calculations and molecular dynamics simulations (with a machine-learned force field) to visualize Pd deposition and migration on the nanoparticle surface, respectively. Our results suggest that the nanoparticle–support interface restricts nanoparticle accessibility to Pd deposition, which consequently controls the Pd ensemble size, illustrating the critical role of nanoscale wetting phenomena during bimetallic catalyst preparation.

36 MATERIALS SCIENCE↗