Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Probabilistic Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Mass Detection for Heavy-Duty Vehicles using Gaussian Belief Propagation

Predicting vehicle mass is critical to accurately estimate energy use and emissions of commercial trucks. However, data from vehicle telematics is often not at sufficient temporal resolution or accuracy for use in model-based detection methods. In this work, a new statistical mass prediction technique is described for heavy-duty vehicles that incorporates the use Gaussian Belief Propagation (GBP) for probabilistic inference. Similar to Bayesian inference models, the GBP model typically requires less labeled training data than other contemporary machine learning techniques. First, a factor graph is constructed, and a set of Gaussian belief nodes with associated means and variances are fitted to the training data. To better handle noisy input data, the GBP mass prediction model utilizes a k-nearest factors (kNF) algorithm for probabilistic inference on unseen testing data. The proposed method is compared with a classical weighted k-nearest neighbors (kNN) regressor. This statistical kNF-GBP model works even with low-quantity, low-quality initial training data, while being capable of realtime mass estimation. Unlike the kNN regressor, the GBP model produces a measure of uncertainty with its predictions. The proposed method is validated using curve-sampled driving data collected from multiple cloud-connected Class 8 regional haul diesel trucks. Both the kNN regressor and the kNF-GBP mass prediction model were able to predict payload mass with coefficients of determination above 0.97 with minimal data preprocessing.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Probabilistic Design Methodology and its Application to the Design of an Umbilical Retract Mechanism

A lot has been learned from past experience with structural and machine element failures. The understanding of failure modes and the application of an appropriate design analysis method can lead to improved structural and machine element safety as well as serviceability. To apply Probabilistic Design Methodology (PDM), all uncertainties are modeled as random variables with selected distribution types, means, and standard deviations. It is quite difficult to achieve a robust design without considering the randomness of the design parameters which is the case in the use of the Deterministic Design Approach. The US Navy has a fleet of submarine-launched ballistic missiles. An umbilical plug joins the missile to the submarine in order to provide electrical and cooling water connections. As the missile leaves the submarine, an umbilical retract mechanism retracts the umbilical plug clear of the advancing missile after disengagement during launch and retrains the plug in the retracted position. The design of the current retract mechanism in use was based on the deterministic approach which puts emphasis on factor of safety. A new umbilical retract mechanism that is simpler in design, lighter in weight, more reliable, easier to adjust, and more cost effective has become desirable since this will increase the performance and efficiency of the system. This paper reports on a recent project performed at Tennessee State University for the US Navy that involved the application of PDM to the design of an umbilical retract mechanism. This paper demonstrates how the use of PDM lead to the minimization of weight and cost, and the maximization of reliability and performance.

Onyebueke, Landon↗

Incorporating polar field data for improved solar flare prediction

In this paper, we consider incorporating data associated with the sun’s north and south polar field strengths to improve solar flare prediction performance using machine learning models. When used to supplement local data from active regions on the photospheric magnetic field of the sun, the polar field data provides global information to the predictor. While such global features have been previously proposed for predicting the next solar cycle’s intensity, in this paper we propose using them to help classify individual solar flares. We conduct experiments using HMI data employing four different machine learning algorithms that can exploit polar field information. Additionally, we propose a novel probabilistic mixture of experts model that can simply and effectively incorporate polar field data and provide on-par prediction performance with state-of-the-art solar flare prediction algorithms such as the Recurrent Neural Network (RNN). Our experimental results indicate the usefulness of the polar field data for solar flare prediction, which can improve Heidke Skill Score (HSS2) by as much as 10.1%.

79 ASTRONOMY AND ASTROPHYSICS↗

Decentralized Collaborative Learning with Probabilistic Data Protection

We discuss future directions of Blockchain as a collaborative value co-creation platform, in which network participants can gain extra insights that cannot be accessed when disconnected from the others. As such, we propose a decentralized machine learning framework that is carefully designed to respect the values of democracy, diversity, and privacy. Specifically, we propose a federated multi-task learning framework that integrates a privacy-preserving dynamic consensus algorithm. We show that a specific network topology called the expander graph dramatically improves the scalability of global consensus building. We conclude the paper by making some remarks on open problems.

Ide, Tsuyoshi↗

Scalable statistical inference of photometric redshift via data subsampling

Handling big data has largely been a major bottleneck in traditional statistical models. Consequently, when accurate point prediction is the primary target, machine learning models are often preferred over their statistical counterparts for bigger problems. But full probabilistic statistical models often outperform other models in quantifying uncertainties associated with model predictions. We develop a data-driven statistical modeling framework that combines the uncertainties from an ensemble of statistical models learned on smaller subsets of data carefully chosen to account for imbalances in the input space. We demonstrate this method on a photometric redshift estimation problem in cosmology, which seeks to infer a distribution of the redshift—the stretching effect in observing the light of far-away galaxies—given multivariate color information observed for an object in the sky. Our proposed method performs balanced partitioning, graph-based data subsampling across the partitions, and training of an ensemble of Gaussian process models.

data subsampling↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

How Can Probabilistic Solar Power Forecasts Be Used to Lower Costs and Improve Reliability in Power Spot Markets? A Review and Application to Flexiramp Requirements

Net load uncertainty in electricity spot markets is rapidly growing. There are five general approaches by which system operators and market participants can use probabilistic forecasts of wind, solar, and load to help manage this uncertainty. These include operator situation awareness, resource risk hedging, reserves procurement, definition of contingencies, and explicit stochastic optimization. We review these approaches, and then provide a case study in which a method for using probabilistic solar forecasts to define needs for reserves is developed and evaluated. The case study has three parts. First, we describe building blocks for enhancing the Watt-Sun solar forecasting system to produce probabilistic irradiance and power forecasts. Second, relationships between Watt-Sun forecasts for multiple sites in California and the system's need for flexible ramp capability (flexiramp) are defined by machine learning and statistical methods. Third, the performance of present methods to defining flexiramp requirements, which are not conditioned on weather and renewables forecasts, is compared with that of probabilistic solar forecast-based requirements, using a multi-timescale production costing model with an 1820-bus representation of the WECC power system. Significant potential savings in fuel and flexiramp procurement costs from using solar-informed reserve requirements are found.

14 SOLAR ENERGY↗

Forecasting Marine Sediment Properties with Geospatial Machine Learning

Using a combination of geospatial machine learning prediction and sediment thermodynamic/physical modeling, we have developed a novel software workflow to create probabilistic maps of geoacoustic and geomechanical sediment properties of the global seabed. This new technique for producing reliable estimates of seafloor properties can better support Naval operations relying on sonar performance and seabed strength, can constrain models of shallow tomographic structure important for nuclear treaty compliance monitoring/detection, and can provide constraints on the distribution and inventory of shallow methane gas and gas hydrate accumulations on the continental shelves.

58 GEOSCIENCES↗

Quantifying the Known Unknown: Including Marine Sources of Greenhouse Gases in Climate Modeling

Researchers have recently estimated that Arctic submarine permafrost currently traps 60 billion tons of methane and contains 560 billion tons of organic carbon in seafloor sediments and soil, a giant pool of carbon with potentially large feedbacks on the climate system. Unlike terrestrial permafrost, the submarine permafrost system has remained a “known unknown” because of the difficulty in acquiring samples and measurements. Consequently, this potentially large carbon stock never yet considered in global climate models or policy discussions, represents a real wildcard in our understanding of Earth’s climate. This report summarizes our group’s effort at developing a numerical modeling framework designed to produce a first-of-its-kind estimate of Arctic methane gas releases from the marine sediments to the water column, and potentially to the atmosphere, where positive climate feedback may occur. Newly developed modeling capability supported by the Laboratory Directed Research and Development (LDRD) program at Sandia National Laboratories now gives us the ability to probabilistically map gas distribution and quantity in the seabed by using a hybrid approach of geospatial machine learning, and predictive numerical thermodynamic ensemble modeling. The novelty in this approach is its ability to produce maps of useful data in regions that are only sparsely sampled, a common challenge in the Arctic, and a major obstacle to progress in the past. By applying this model to the circum-Arctic continental shelves and integrating the flux of free gas from in situ methanogenesis and dissociating gas hydrates from the sediment column under climate forcing, we can provide the most reliable estimate of a spatially and temporally varying source term for greenhouse gas flux that can be used by global oceanographic circulation and Earth system models (such as DOE’s E3SM). The result will allow us to finally tackle the wildcard of the submarine permafrost carbon system, and better inform us about the severity of future national security threats that sustained climate change poses.

54 ENVIRONMENTAL SCIENCES↗

Coordinated Ramping Product and Regulation Reserve Procurements in CAISO and MISO using Multi-Scale Probabilistic Solar Power Forecasts (Pro2R)

How can probabilistic solar forecasts lower costs and improve reliability for independent system operator (ISO) markets? We tackle this question in three steps. First, we enhance an existing solar forecasting system to provide well-calibrated hours-ahead probabilistic forecasts. We then relate the degree of uncertainty in those forecasts to error distributions for net load ramps for the California ISO (CAISO) using statistical and machine learning methods. Projected net load errors conditioned on solar uncertainty are translated into flexible ramp requirements that therefore reflect real-time meteorological and solar conditions, improving on typical ISO procedures. Finally, a multi-period look-ahead production cost model quantifies how conditional ramp requirements can a) decrease operating costs by lowering requirements compared to often conservative unconditional methods, and b) reduce generation scarcity events and consequently improve reliability by increasing flexibility requirements at times when unconditional forecast-based requirements understate actual ramp uncertainty. In addition to the products just described (quantification of solar uncertainty, its translation into requirements for ramp capability product, and quantification of the benefits of more accurate ramp requirements), this project also developed a visualization system that alerts system operators of ramp and uncertainty conditions within the network based on solar forecasts. The system is called Resource Forecast and Ramp Visualization for Situational Awareness (RaVIS). These four products represent significant advances in the state-of-the-art of probabilistic solar forecasting, development of weather-informed reserve requirements, production costing methods for estimating the benefits of more accurate reserve requirements, and visualization of system status, respectively. Yet the products are also practical and can be immediately implemented, potentially enabling system operators to save millions of dollars in ramp product procurement costs per year.

14 SOLAR ENERGY↗

Active learning using hybrid surrogate tool life modeling for machining process optimization

Here, this paper describes an active learning approach for part-to-part iterative machining process optimization using a hybrid surrogate tool life model. A probabilistic interpolating tool life model is developed by combining the empirical Taylor-type tool life equation and the model fit error. The probabilistic tool life model is then used to calculate the machining cost per part distribution. The optimal machining parameters are selected using an expected improvement in machining cost per part criterion. The method is validated numerically using experimental results; the results show a median convergence error of 2.2% after three tests over 400 simulations. The method is validated experimentally on two industrial applications for Ti-6Al-4V roughing resulting in a cost per part reduction greater than 23% after two tests. The described method is a robust solution for rapid convergence to optimal machining parameters in an industrial production environment.

Active learning↗

Probabilistic Day-Ahead Forecasting Using an Analog Ensemble Approach for Wind Farm Grid Services

Wind resource assessment and wind power forecasting are used in research and industry to anticipate future power output at scales ranging from individual wind turbines to entire wind farms. Probabilistic day-ahead wind forecasting is useful for anticipating how a wind farm could potentially participate in the day-ahead market by providing upper and lower bounds for expected power generation, thus informing grid operators of its uncertainty. Understanding this uncertainty is part of a larger project focused on building a platform that combines efforts in weather forecasting, aerodynamic and economic modeling to create maximum value of a wind plant to better provide services to the grid. This effort is also known as the Atmosphere to Electrons to Grid (A2E2G) project. One method for producing a probabilistic forecast is through the analog ensemble approach (Delle Monache et al., 2011). This method leverages historical forecasts and their corresponding observations as a training data set from which future forecasts can be made. For some future forecast, the most similar historical forecasts (analogs) are identified on a regular time basis such as once per a 3-hour window. The most similar analogs, based on a metric such as root mean square error (RMSE), are recorded and their corresponding verifying observations are used as an ensemble member for this future forecast. Prior work in this area demonstrates improvements over raw Numerical Weather Prediction (NWP) forecasts and shows skill similar to techniques such as logistic regression and machine learning (Delle Monache et al., 2013; Alessandrini et al., 2015). Here, we take the High-Resolution Rapid Refresh model (HRRR) day-ahead forecast (0-36 hours) to create a probabilistic day-ahead forecast using an analog ensemble approach. The HRRR has an hourly temporal resolution, with a spatial resolution of 3 km. The 12 UTC HRRR model run is downloaded every day for one year from August 2019 - July 2020, with the first 11 months serving as a bank of analogs from which the forecasting algorithm can create a probabilistic forecast. Once downloaded, the original HRRR forecast is temporally interpolated to 5-minutes, aligning with both the temporal resolution of the observations as well as the timescale relevant for day-ahead power forecasts. The forecast is validated at the M2 tower at the Flatirons Campus of the National Renewable Energy Laboratory (NREL) at a typical wind turbine height of 80 m. Variables such as wind speed, wind direction, and turbulence intensity are incorporated into the probabilistic forecast model and weighted according to their relative importance to the forecast. Based on metrics such as mean bias error (MBE), mean absolute error (MAE), and root mean square error, the analog ensemble forecast outperforms the raw HRRR forecast during the testing period of July 2020. Figure 1 illustrates an example day-ahead forecast compared against the verifying observations. The general variability and ramps are captured throughout the day, with potential to further improve the analog ensemble model through machine learning techniques.

numerical weather prediction↗

Review of Onsite Temperature and Solar Forecasting Models to Enable Better Building Design and Operations

Advanced building controls and energy optimization for new constructions and retrofits rely on accurate weather data. Traditionally, most studies utilize airport weather information as the decision inputs. However, most buildings are in environments that are quite different than those at the airport miles away. Tree cover, adjacent buildings, and micro-climate effects caused by the larger surrounding area can all yield deviations in air temperature, humidity, solar irradiance, and wind that are large enough to influence design and operation decisions. In order to overcome this challenge, there are many prior studies on developing weather forecasting algorithms from micro-to meso-scales. Additionally, this paper reviews and complies knowledge on common weather data resources, data processing methodologies and forecasting techniques of weather information. Commonly used statistical, machine learning and physical-based models are discussed and presented as two major categories: deterministic forecasting and probabilistic forecasting. Finally, evaluation metrics for forecasting errors are listed and discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Twisty-puzzle-inspired approach to Clifford synthesis

The problem of decomposing an arbitrary Clifford element into a sequence of Clifford gates is known as Clifford synthesis. Drawing inspiration from similarities between this and the famous Rubik's cube twisty puzzle, here we develop a machine learning approach for Clifford synthesis based on learning an approximation to the distance to the identity. This approach is probabilistic and computationally intensive. However, when a decomposition is successfully found, it often involves fewer gates than the decomposition methods used in the Qiskit decomposition protocol, which uses a combination of several well-known Clifford decomposition schemes. Additionally, our approach is much more flexible than existing algorithms in that arbitrary gate sets, device topologies, and gate fidelities may be incorporated, thus allowing for the approach to be tailored to a specific device.

97 MATHEMATICS AND COMPUTING↗

Randomized Algorithms for Scientific Computing (RASC)

Randomized algorithms have propelled advances in artificial intelligence (AI) and represent a foundational research area in advancing AI for Science. Future advancements in DOE Office of Science priority areas such as climate science, astrophysics, fusion, advanced materials, combustion, and quantum computing all require randomized algorithms for surmounting challenges of complexity, robustness, and scalability. Advances in data collection and numerical simulation have changed the dynamics of scientific research and motivate the need for randomized algorithms. For instance, advances in imaging technologies such as X-ray ptychography, electron microscopy, electron energy loss spectroscopy, or adaptive optics lattice light-sheet microscopy collect hyperspectral imaging and scattering data in terabytes, at breakneck speed enabled by state-of-the-art detectors. The data collection is exceptionally fast compared with its analysis. Likewise, advances in high-performance architectures have made exascale computing a reality and changed the economies of scientific computing in the process. Floating-point operations that create data are essentially free in comparison with data movement. Thus far, most approaches have focused on creating faster hardware. Ironically, this faster hardware has exacerbated the problem by making data still easier to create. Under such an onslaught, scientists often resort to heuristic deterministic sampling schemes (e.g., low-precision arithmetic, sampling every nth element) and sacrifice potentially valuable accuracy. Dramatically better results can be achieved via randomized algorithms, reducing the data size as much as or more than naive deterministic subsampling can achieve, while retaining the high accuracy of computing on the full data set. By randomized algorithms we mean those algorithms that employ some form of randomness in internal algorithmic decisions to accelerate time to solution, increase scalability, or improve reliability. Examples include matrix sketching for solving large-scale least-squares problems (see Figure 1) and stochastic gradient descent for training machine learning models. We are not recommending heuristic methods but rather randomized algorithms that have certificates of correctness and probabilistic guarantees of optimality and near-optimality. Such approaches can be useful beyond acceleration, for example, in understanding how to avoid measure zero worst-case scenarios that plague methods such as QR matrix factorization.

97 MATHEMATICS AND COMPUTING↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]↗

Predicting Solar Energetic Particles Using SDO/HMI Vector Magnetic Data Products and a Bidirectional LSTM Network

Solar energetic particles (SEPs) are an essential source of space radiation, and are hazardous for humans in space, spacecraft, and technology in general. In this paper, we propose a deep-learning method, specifically a bidirectional long short-term memory (biLSTM) network, to predict if an active region (AR) would produce an SEP event given that (i) the AR will produce an M- or X-class flare and a coronal mass ejection (CME) associated with the flare, or (ii) the AR will produce an M- or X-class flare regardless of whether or not the flare is associated with a CME. The data samples used in this study are collected from the Geostationary Operational Environmental Satellite's X-ray flare catalogs provided by the National Centers for Environmental Information. We select M- and X-class flares with identified ARs in the catalogs for the period between 2010 and 2021, and find the associations of flares, CMEs, and SEPs in the Space Weather Database of Notifications, Knowledge, Information during the same period. Each data sample contains physical parameters collected from the Helioseismic and Magnetic Imager on board the Solar Dynamics Observatory. Experimental results based on different performance metrics demonstrate that the proposed biLSTM network is better than related machine-learning algorithms for the two SEP prediction tasks studied here. We also discuss extensions of our approach for probabilistic forecasting and calibration with empirical evaluation

79 ASTRONOMY AND ASTROPHYSICS↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗