Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “probabilistic machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Mesoscale informed parameter estimation through machine learning: A case-study in fracture modeling

Scale bridging is a critical need in computational sciences, where the modeling community has developed accurate physics models from first principles, of processes at lower length and time scales that influence the behavior at the higher scales of interest. However, it is not computationally feasible to incorporate all of the lower length scale physics directly into upscaled models. This is an area where machine learning has shown promise in building emulators of the lower length scale models, which incur a mere fraction of the computational cost of the original higher fidelity models. We demonstrate the use of machine learning using an example in materials science estimating continuum scale parameters by emulating, with uncertainties, complicated mesoscale physics. Additionally, we describe a new framework to emulate the fine scale physics, especially in the presence of microstructures, using machine learning, and showcase its usefulness by providing an example from modeling fracture propagation. Our approach can be thought of as a data-driven dimension reduction technique that yields probabilistic emulators. Our results show well-calibrated predictions for the quantities of interests in a low-strain simulation of fracture propagation at the mesoscale level. Furthermore, on average, we achieve ~10% relative errors on time-varying quantities like total damage and maximum stresses. Successfully replicating mesoscale scale physics within the continuum models is a crucial step towards predictive capability in multi-scale problems.

36 MATERIALS SCIENCE↗

Time-Resolved Line Shapes of Single Quantum Emitters via Machine Learned Photon Correlations

Solid-state single-photon emitters (SPEs) are quantum light sources that combine atomlike optical properties with solid-state integration and fabrication capabilities. SPEs are hindered by spectral diffusion, where the emitter’s surrounding environment induces random energy fluctuations. Timescales of spectral diffusion span nanoseconds to minutes and require probing single emitters to remove ensemble averaging. Photon correlation Fourier spectroscopy (PCFS) can be used to measure time-resolved single emitter line shapes, but is hindered by poor signal-to-noise ratio in the measured correlation functions at early times due to low photon counts. Here, we develop a framework to simulate PCFS correlation functions directly from diffusing spectra that match well with experimental data for single colloidal quantum dots. We use these simulated datasets to train a deep ensemble autoencoder machine learning model that outputs accurate, noiseless, and probabilistic reconstructions of the noisy correlations. Using this model, we obtain reconstructed time-resolved single dot emission line shapes at timescales as low as 10 ns, which are otherwise completely obscured by noise. This enables PCFS to extract optical coherence times on the same timescales as Hong-Ou-Mandel two-photon interference, but with the advantage of providing spectral information in addition to estimates of photon indistinguishability. Further, our machine learning approach is broadly applicable to different photon correlation spectroscopy techniques and SPE systems, offering an enhanced tool for probing single emitter line shapes on previously inaccessible timescales.

74 ATOMIC AND MOLECULAR PHYSICS↗

Machine Learning Derived Dynamic Operating Reserve Requirements in High-Renewable Power Systems

Accurately forecasting wind and solar power output poses challenges for deeply decarbonized electricity systems. Grid operators must commit resources to provide reserves to ensure reliable operations in the face of forecast errors, a process which can increase fuel consumption and emissions. We apply neural network-based machine learning to expand the usefulness of median point forecast data by creating probabilistic distributions of short-term uncertainty in demand, wind, and solar forecasts that adapt to prevailing grid conditions. Machine learning derived estimates of forecast errors compare favorably to estimates based on incumbent methods. Reserves derived from machine learning are usually smaller than values derived using incumbent methods, which enables fuel savings during most hours. Machine learning reserves are generally larger than incumbent reserves during times of higher forecast error, potentially improving system reliability. Performance is tested using multi-stage production simulation modeling of the California Independent System Operator (CAISO) system. Machine learning reserves provide production cost and greenhouse gas (GHG) emission reductions of approximately 0.3% relative to historical 2019 requirements. Savings in the 2030 timeframe are highly dependent on battery storage capacity. At lower levels of battery capacity, savings of 0.4% from machine learning reserves are shown. Significant quantities of battery storage are expected to be added to meet California's resource adequacy needs and GHG reduction targets. Addition of these batteries saturate reserve needs and results in minimal within-hour balancing costs in 2030.

24 POWER TRANSMISSION AND DISTRIBUTION↗

EFIT-Prime: Probabilistic and physics-constrained reduced-order neural network model for equilibrium reconstruction in DIII-D

We introduce EFIT-Prime, a novel machine learning surrogate model for EFIT (Equilibrium FIT) that integrates probabilistic and physics-informed methodologies to overcome typical limitations associated with deterministic and ad hoc neural network architectures. EFIT-Prime utilizes a neural architecture search-based deep ensemble for robust uncertainty quantification, providing scalable and efficient neural architectures that comprehensively quantify both data and model uncertainties. Physically informed by the Grad–Shafranov equation, EFIT-Prime applies a constraint on the current density J tor and a smoothness constraint on the first derivative of the poloidal flux, ensuring physically plausible solutions. Furthermore, the spatial location of the diagnostics is explicitly incorporated in the inputs to account for their spatial correlation. Extensive evaluations demonstrate EFIT-Prime's accuracy and robustness across diverse scenarios, most notably showing good generalization on negative-triangularity discharges that were excluded from training. Timing studies indicate an ensemble inference time of 15 ms for predicting a new equilibrium, offering the possibility of plasma control in real-time, if the model is optimized for speed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

Logical Activation Functions v.1.1

SAND2024-01501O Logical Activation Functions software is a PyTorch implementation from the paper, "Logical Activation Functions for Training Arbitrary Probabilistic Boolean Logic." The activation functions approximate logit-space marginalization of probabilistic truth tables from probabilistic interpretations of inputs. They also provide a general methodology to approximate logical relationships between abstract antecedents and consequents for machine learning architectures. They do not target any specific application or use-case. By training probabilistic truth tables, these activation functions can capture more expressive relationships in a neural network than typical elementwise activation functions. This code is only designed for a single compute node with a GPU and is limited to machine learning architectures than can fit within the memory of a single GPU. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Duersch, Jed↗

Unified architecture for data-driven metadata tagging of building automation systems

This article presents a Unified Architecture (UA) for automated point tagging of Building Automation System (BAS) data, based on a combination of data-driven approaches. Advanced energy analytics applications—including fault detection and diagnostics and supervisory control—have emerged as a significant opportunity for improving the performance of our built environment. Effective application of these analytics depends on harnessing structured data from the various building control and monitoring systems, but typical BAS implementations do not employ any standardized metadata schema. While standards such as Project Haystack and Brick Schema have been developed to address this issue, the process of structuring the data, i.e., tagging the points to apply a standard metadata schema, has, to date, been a manual process. This process is typically costly, labor-intensive, and error-prone. In this work we address this gap by proposing a UA that automates the process of point tagging by leveraging the data accessible through connection to the BAS, including time-series data and the raw point names. The UA intertwines supervised classification and unsupervised clustering techniques from machine learning and leverages both their deterministic and probabilistic outputs to inform the point tagging process. Furthermore, we extend the UA to embed additional input and output data-processing modules that are designed to address the challenges associated with the real-time deployment of this automation solution. We test the UA on two datasets for real-life buildings: (i) commercial retail buildings and (ii) office buildings from the National Renewable Energy Laboratory (NREL) campus. We report the proposed methodology correctly applied 85–90% and 70–75% of the tags in each of these test scenarios, respectively for two significantly different building types used for testing UA's fully-functional prototype. The proposed UA, therefore, offers promising approach for automatically tagging BAS data as it reaches close to 90% accuracy. Further building upon this framework to algorithmically identify the equipment type and their relationships is an apt future research direction to pursue.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Systems and methods for improved aircraft safety

Methods and system for improved aircraft safety are described herein. A machine learning model may be trained using historical flight data for a number of previously flown flights where an automation function autonomously disengaged. The trained machine learning model may be used to generate probabilistic alert rules. The probabilistic alert rules may be used by a computing device onboard an aircraft to provide contextual information to flight crew relating to engagement status for one or more automation functions of the aircraft.

Sherry, Lance↗

Applying AI/ML Techniques to U.S. Nuclear Operating Experience Program

Idaho National Laboratory (INL) has provided technical assistance to the U.S. Nuclear Regulatory Commission (NRC) in reliability and risk analysis including the operating experience (OpE) program since the 1980s. The U.S. nuclear OpE program provides input parameters to the NRC Standardized Plant Analysis Risk models and the industry probabilistic risk assessment (PRA) models. While earlier PRA focuses were on at-power, internal event analysis, the risks from external hazards and during low power shutdown (LPSD) operation could be significant and the needs to develop LPSD PRA and external hazards PRA are on the rise. One issue in developing LPSD PRA is the reasonable estimation of shutdown initiative event (SDIE) frequencies. INL has developed and is maintaining an SDIE database for the NRC. However, this database is based on the reviewing of Licensee Event Reports (LERs), which is believed to be only a subset of “actual” shutdown initiating events occurred in the industry. This paper investigates a new approach to identify and characterize shutdown initiating events from the Institute of Nuclear Power Operations (INPO) industry database using machine learning techniques. The main process in this approach is to find out the relationship between key words in event descriptions and the SDIE categories as in the NRC SDIE database. The relationship can then be applied to the INPO database and search for SDIEs.

99 GENERAL AND MISCELLANEOUS↗

Operational Probabilistic Tools for Solar Uncertainty (OPTSUN) (Final Project Report for DOE Solar Forecasting II Project)

Increasing levels of solar PV can challenge system operations and may require novel methods to operate the power system reliably and efficiently. Power system operating plans generally use deterministic forecasts, in which the variable energy resources are represented by the expected value for each interval of the decision horizon. Probabilistic forecasts are relatively new but have the potential to address the shortfalls of deterministic forecasts. However, understanding how best to use such forecasts is still a key gap in industry and was the focus of this project. The project had three workstreams. In a forecasting workstream, improvements were made to baseline probabilistic forecasts using a number of new approaches such as machine learning methods and improved input data. In a design workstream, advanced simulation tools used these forecasts to investigate newly proposed reserve determination methods. Lastly, in a demonstration workstream a scheduling management platform (SMP) was developed to leverage probabilistic forecasts in a modular and customizable manner. In order to study the benefits that could be accrued, the project team collaborated with three utility partners (Duke Energy, Southern Company and Hawaiian Electric) to deliver improved probabilistic forecasts for each region and to model each region in case studies using advanced production cost modeling tools. Different methods to determine operating reserve requirements from probabilistic forecasts were developed, simulated, and tested across each region. The benefits of using these newly proposed methods varied by utility, but, in general, using probabilistic forecasts as well as historical data to set the reserve requirements seems to improve reliability related results, with less risk of reserve or supply shortfalls. The cost implications were not always straightforward; in some cases the new methods could show a reduction in expected operating costs, but often the increase in reserves associated with better risk mitigation using probabilistic forecasts could result in an increase in operating costs in the simulations. The SMP tool was developed to process probabilistic forecasts from their initial receipt through to scheduling decisions. This open-source tool consists of several modules for scenario development, reserve requirements calculation, and visualization. The SMP tool was demonstrated to a wide range of operators and stakeholders at all three utilities and further improved based on their feedback. The tool will be available on www.epri.com/optsun. The proposed probabilistic information-based reserve determination approaches have the potential to be implemented by different regions to ensure an economic and reliable power system operation on power systems integrating increasing levels of variable renewable resources. The innovative yet practical methods developed in this project demonstrated tangible benefits from using probabilistic forecasts beyond just study-based assessments to include three unique balancing areas. The demonstrated benefits across the multiple utility environments, are expected to provide system operators in all regions the confidence required and a platform to adopt the new forecasting and operating methods.

14 SOLAR ENERGY↗

Predicting peak day and peak hour of electricity demand with ensemble machine learning

Battery energy storage systems can be used for peak demand reduction in power systems, leading to significant economic benefits. Two practical challenges are 1) accurately determining the peak load days and hours and 2) quantifying and reducing uncertainties associated with the forecast in probabilistic risk measures for dispatch decision-making. In this study, we develop a supervised machine learning approach to generate 1) the probability of the next operation day containing the peak hour of the month and 2) the probability of an hour to be the peak hour of the day. Guidance is provided on preparation and augmentation of data as well as selection of machine learning models and decision-making thresholds. The proposed approach is applied to the Duke Energy Progress system and successfully captures 69 peak days out of 72 testing months with a 3% exceedance probability threshold. On 90% of the peak days, the actual peak hour is among the 2 h with the highest probabilities.

25 ENERGY STORAGE↗

Risk-Informed Condition Evaluation of Solar-centered Energy Generation and Distribution Networks through Bayesian Learning and Inference

We develop a methodology based on Bayesian inference over Probabilistic Graphical Models (PGMs) to understand and quantify risk in solar-centered grids using targeted measurements and learned system behavior. Being non-prescriptive but, rather, able to infer system behavior and, ultimately, address risk queries from data, our machine learning-type paradigm is tailored for diverse topologies and threat scenarios often associated with distributed energy generation and photovoltaic distributed energy resources (PV-DERs) in particular. We describe algorithmic processes for: (i) learning the structure of PGMs that result from attack-prone PV-DER-proliferated distribution systems, (ii) quantifying cause-effect relationships, and (iii) evaluating risk queries based on diverse evidence. The contributions are illustrated on a residential grid subject to output impairment attacks on its PV-DER infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Probabilistic neural networks for fluid flow surrogate modeling and data recovery

We consider the use of probabilistic neural networks for fluid flow surrogate modeling and data recovery. This framework is constructed by assuming that the target variables are sampled from a Gaussian distribution conditioned on the inputs. Consequently, the overall formulation sets up a procedure to predict the hyperparameters of this distribution which are then used to compute an objective function given training data. We demonstrate that this framework has the ability to provide for prediction confidence intervals based on the assumption of a probabilistic posterior, given an appropriate model architecture and adequate training data. The applicability of the present framework to cases with noisy measurements and limited observations is also assessed. To demonstrate the capabilities of this framework, we consider canonical regression problems of fluid dynamics from the viewpoint of reduced-order modeling and spatial data recovery for four canonical data sets. The examples considered in this study arise from (i) the shallow-water equations, (ii) a two-dimensional cylinder flow, (iii) the wake of a NACA0012 airfoil with a Gurney flap, and (iv) the NOAA sea surface temperature data set. Furthermore, the present results indicate that the probabilistic neural network not only produces a machine-learning-based fluid flow surrogate model but also systematically quantifies the uncertainty therein to assist with model interpretability.

42 ENGINEERING↗

Accurate and Fast Anomaly Detection in Additive Composite-Based Manufacturing using Thermal Cameras

Today, large-scale additive manufacturing with plastics and composite materials requires continuous monitoring by experienced staff to prevent, detect and correct anomalous events affecting the performance of the printed part. We address the complexity of this demanding task by designing a camera-based anomaly detection system utilizing probabilistic principal component analysis (PPCA). This is a machine learning technique is trained with thermal images collected during normal operation of the large-scale printer (Cincinnati BAAM). This technique is advantageous for practical applications as there is no need to artificially introduce anomalous conditions into model training. During deployment, we challenge this model by introducing deliberate variations of the extruder speed. We reduce extrusion speed to a lower level, between 70 and 95% of the nominal value to collected test images. Our results show that images are easily identified as anomalous for extruder speeds at or below 85% of the nominal speed, meaning that an anomalous reduction of the material deposition rate can be detected within seconds of its onset. We show that our results are robust to (a) camera-to-camera variability and (b) print-to-print variability.

Pike, John [ORNL]↗

Denoising diffusion probabilistic models for generative alloy design

Inverse material design is an extremely challenging optimization task made difficult by, in part, the highly nonlinear relationship linking performance with composition. Quantitative approaches have improved significantly owing to advances in high throughput experimentation and computational thermodynamics. However, existing physics-based tools are mostly forward models; input a chemistry and obtain a prediction. More recently the materials community has leveraged advances in the machine learning community to establish novel inverse design frameworks. Very recently denoising diffusion probabilistic models have been shown to be extremely powerful generators producing synthetic data of various modalities e.g. images, text, audio, tables, etc.. In this work a novel framework for alloy design and optimization is proposed leveraging these class of models. Five key generative tasks are demonstrated (1) unconditional generation (2) composition conditioned generation (3) property conditioned generation (4) multi-feedstock conditioned generation and (5) generative optimization. These methods were tested on three case studies: high entropy alloy design, superalloy binder jet additive manufacturing, and in-situ dual-feedstock wire-arc additive manufacturing. Results indicate that the established models are extremely flexible, expressive, and robust. The architecture’s flexibility and training procedure empower the model to learn complex intra-compositional and composition-property relationships. Furthermore, the probabilistic nature of these models makes them well suited for addressing solution non-uniqueness and tackling uncertainty quantification tasks. While the fidelity and quantity of the underlying training data is paramount, we envision that future alloy design frameworks will make extensive use of these kinds of machine learning models as “search” tools bolstering the utility of experimental and computational approaches.

36 MATERIALS SCIENCE↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments

There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.

97 MATHEMATICS AND COMPUTING↗

A Bayesian Deep Learning Approach to Near-Term Climate Prediction

Since model bias and associated initialization shock are serious shortcomings that reduce prediction skills in state-of-the-art decadal climate prediction efforts, we pursue a complementary machine-learning-based approach to climate prediction. The example problem setting we consider consists of predicting natural variability of the North Atlantic sea surface temperature on the interannual timescale in the pre-industrial control simulation of the Community Earth System Model. While previous works have considered the use of recurrent networks such as convolutional LSTMs and reservoir computing networks in this and other similar problem settings, we currently focus on the use of feedforward convolutional networks. In particular, we find that a feedforward convolutional network with a Densenet architecture is able to outperform a convolutional LSTM in terms of predictive skill. Next, we go on to consider a probabilistic formulation of the same network based on Stein variational gradient descent and find that in addition to providing useful measures of predictive uncertainty, the probabilistic (Bayesian) version improves on its deterministic counterpart in terms of predictive skill. Finally, we characterize the reliability of the ensemble of machine learning models obtained in the probabilistic setting by using analysis tools developed in the context of ensemble numerical weather prediction.

54 ENVIRONMENTAL SCIENCES↗