Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Probabilistic Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Optimal Transport as a Tool for Scientific Discovery in Radiation Biology

This report summarizes findings from research conducted for the “Exploration of the Poten tial for Artificial Intelligence and Machine Learning to Advance Low-Dose Radiation Biology Re search” (RadBio-AI) program, supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, under Awards KP1601011/FWP CC121 and KP1601017/FWP CC121. The research reported here was undertaken in an effort to assess the potential of optimal measure transport methods as components within the larger scope of a com putational framework envisioned to support research in the radiation biology domain. Within this effort, our interest centered on enabling a unified generic framework where probabilistic modeling, inference, and statistical learning can be carried out for a wide range of data distributions. As described next in Section 1 (and in more detail in our original publication), optimal measure transport offers the possibility of such unified approach.

97 MATHEMATICS AND COMPUTING↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Electrocardiographic changes predate Parkinson’s disease onset

Autonomic nervous system involvement precedes the motor features of Parkinson’s disease (PD). Our goal was to develop a proof-of-concept model for identifying subjects at high risk of developing PD by analysis of cardiac electrical activity. We used standard 10-s electrocardiogram (ECG) recordings of 60 subjects from the Honolulu Asia Aging Study including 10 with prevalent PD, 25 with prodromal PD, and 25 controls who never developed PD. Various methods were implemented to extract features from ECGs including simple heart rate variability (HRV) metrics, commonly used signal processing methods, and a Probabilistic Symbolic Pattern Recognition (PSPR) method. Extracted features were analyzed via stepwise logistic regression to distinguish between prodromal cases and controls. Stepwise logistic regression selected four features from PSPR as predictors of PD. The final regression model built on the entire dataset provided an area under receiver operating characteristics curve (AUC) with 95% confidence interval of 0.90 [0.80, 0.99]. The five-fold cross-validation process produced an average AUC of 0.835 [0.831, 0.839]. We conclude that cardiac electrical activity provides important information about the likelihood of future PD not captured by classical HRV metrics. Machine learning applied to ECGs may help identify subjects at high risk of having prodromal PD.

59 BASIC BIOLOGICAL SCIENCES↗

Learning from learning machines: improving the predictive power of energy-water-land nexus models with insights from complex measured and simulated data

Focal Area(s): Insights gleaned from complex data (observed and simulated) using AI; Predictive modeling through the use of AI techniques and AI-derived model components, including physics- and knowledge-informed models; Energy-water-land nexus and integrated energy systems – models of MultiSector dynamics. Science Challenge: Scientific communities are in need of tools for the computational integration of physics-based models, experimental data, and empirical/observational studies across a broad range of temporal and spatial scales to explore and meet EESSD grand challenges. We require AI algorithms for the discovery of process-drivers in the Earth-energy-human system and to link them with complex measured and simulated data. We aim to understand the energy-water-land nexus, particularly under extreme forcing scenarios and rare events, leveraging scale-aware AI process models, including probabilistic uncertainties, and benefiting from the emerging 5G-enabled landscape-scale sensing and edge computing capabilities. These models are needed to identify instabilities and tipping points that manifest extreme system behaviors with consequences for integrated energy systems and the environment. This work requires fundamental advances in uncertainty quantification, in particular, to identify and model unlikely but catastrophic outliers. Specifically, we need models that are interpretable to domain scientists and, essentially, explainable to public and private stake holders.

54 ENVIRONMENTAL SCIENCES↗

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction↗

A Probabilistic Approach To Selecting Build Configurations in Package Managers

In the past decade software has grown significantly in complexity and scale. Likewise the number of dependencies for most software has increased with typical software packages depending on tens to hundreds of other packages. Such large numbers of dependencies place a significant burden on package managers to correctly maintain dependency lists and constraints between them. Due to this package managers have incorporated sophisticated tooling such as SAT solvers into their dependency management mechanisms. Despite these tools package managers still rely on many human annotated constraints for dependency and version selection. These are error-prone and require a significant amount of labor to constantly update and test. In this paper we propose a methodology to make use of historical build results in selecting the version for package dependencies. Our method utilizes the flexibility of the Spack package manager’s heavily parameterized package configurations to incorporate a machine learning model trained to predict the probability of build outcomes. This work is able to build and install packages with a 13% higher success rate than the default version selection mechanism in Spack.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification for Dissimilar Material Joints Under Corrosion Environment

Abstract Self-Piercing Riveting (SPR) is one of the most commonly used methods for joining dissimilar materials in the automotive industry. These joints are popular due to their adaptability, high performance and short cycle time. However, since these joints involve two dissimilar materials, they are susceptible to galvanic corrosion in the presence of an electrolyte which is driven by the difference in the equilibrium potential of the metals. This can affect the safety and resilience of these joints. In this paper, we focus on galvanic corrosion in Al-Fe SPR joints. A Machine learning (ML) based surrogate model, which is based off of FE simulations, for statistical corrosion analysis is developed. This model enables the resilience and reliability analysis of SPR joints under corrosion environment. In this study, first a physics-based finite element (FE) corrosion model has been developed to simulate the galvanic corrosion between a Fe cathode and an Al anode of a SPR joint. This model takes into account the effect of the crystal microstructure of the Al anode and the precipitation of the corrosion product. Several geometric and environmental factors including crevice gap, roughness of anode, conductivity, pH and the temperature of the electrolyte that effect corrosion are investigated. A thorough Uncertainty Quantification (UQ) analysis is conducted for the overall corrosion behavior of the Fe-Al SPR joints using a novelistic Probabilistic Confidence-Based Adaptive Sampling (PCAS) technique. PCAS is used to train the surrogate model by identifying the critical sampling points and thus reducing the overall computational costs. It is found that the electrolyte temperature has the largest effects on the material loss and needs to be managed closely for better corrosion control. By understanding the corrosion performance and resultant uncertainty impact on SPR joints, the reliability and resilience of these joints can be improved.

36 MATERIALS SCIENCE↗

Machine learning with knowledge constraints for process optimization of open-air perovskite solar cell manufacturing

Perovskite photovoltaics (PV) have achieved rapid development in the past decade in terms of power conversion efficiency of small-area lab-scale devices; however, successful commercialization still requires further development of low-cost, scalable, and high-throughput manufacturing techniques. One of the critical challenges of developing a new fabrication technique is the high-dimensional parameter space for optimization, but machine learning (ML) can readily be used to accelerate perovskite PV scaling. Herein, we present an ML-guided framework of sequential learning for manufacturing process optimization. We apply our methodology to the Rapid Spray Plasma Processing (RSPP) technique for perovskite thin films in ambient conditions. With a limited experimental budget of screening 100 process conditions, we demonstrated an efficiency improvement to 18.5% as the best-in-our-lab device fabricated by RSPP, and we also experimentally found 10 unique process conditions to produce the top-performing devices of more than 17% efficiency, which is 5 times higher rate of success than the control experiments with pseudo-random Latin hypercube sampling. Our model is enabled by three innovations: (a) flexible knowledge transfer between experimental processes by incorporating data from prior experimental data as a probabilistic constraint; (b) incorporation of both subjective human observations and ML insights when selecting next experiments; (c) adaptive strategy of locating the region of interest using Bayesian optimization first, and then conducting local exploration for high-efficiency devices. Furthermore, in virtual benchmarking, our framework achieves faster improvements with limited experimental budgets than traditional design-of-experiments methods (e.g., one-variable-at-a-time sampling). This framework shows the capability of incorporating researchers’ domain knowledge into the ML-guided optimization loop; therefore, it has the potential to facilitate the wider adoption of ML in scaling to perovskite PV manufacturing.

14 SOLAR ENERGY↗

A Framework for Inverse Prediction Using Functional Response Data

Inverse prediction models have commonly been developed to handle scalar data from physical experiments. However, it is not uncommon for data to be collected in functional form. When data are collected in functional form, it must be aggregated to fit the form of traditional methods, which often results in a loss of information. For expensive experiments, this loss of information can be costly. In this study, we introduce the functional inverse prediction (FIP) framework, a general approach which uses the full information in functional response data to provide inverse predictions with probabilistic prediction uncertainties obtained with the bootstrap. The FIP framework is a general methodology that can be modified by practitioners to accommodate many different applications and types of data. We demonstrate the framework, highlighting points of flexibility, with a simulation example and applications to weather data and to nuclear forensics. Results show how functional models can improve the accuracy and precision of predictions.

42 ENGINEERING↗

A hybrid data-driven and model-based approach for computationally efficient stochastic unit commitment and economic dispatch under wind and solar uncertainty

Stochastic unit commitment (UC) and economic dispatch (ED) are imperative in dealing with uncertainty in renewable forecast for power system operation and planning such that the overall expected production cost is minimized over the planning horizon. However, accurate calculation of the expected production cost requires assessment of a very large number of different scenarios of uncertain renewable resources, such as solar and wind, which is practically infeasible to simulate in real time. This article proposes a hybrid datadriven and physics-based model-predictive paradigm to efficiently solve for stochastic unit commitment and economic dispatch considering uncertainty in wind and solar power forecasts. Here, the novelty of the approach lies in decoupling the production cost estimation from the unit commitment and economic dispatch optimization problems under uncertainty without compromising on the fidelity of the solutions. A data-driven machine learning model is first developed to predict the mean optimal production cost. A physics-based inverse problem is then solved to get the stochastic UC and ED profiles from the expected cost. The presented approach considers, for the first time, solar uncertainty in UC/ED determination and enables efficient and accurate propagation of wind and solar uncertainty to estimate the statistics of the production cost. The effectiveness of the developed approach is demonstrated systematically on a stylized RTS-GMLC single-node system. The overall framework predicts the expected cost 62.5% more accurately than the existing state-of-the-art, on unforeseen days during the entire year, and yields, for the first time, the associated physically consistent UC and ED profiles. The solutions are also shown to be flexible in providing adequate daily reserves to address any statistical deviations from probabilistic power forecasts. The computational time associated with the presented method is only about 10 s compared to over 24 h needed for a conventional stochastic UC/ED determination under uncertainty on an Intel Core i9 processor with 32 GB of RAM.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods↗

Stochastic Thermo-Hydro Modeling and Neural Network Surrogate Development for Thermal Resource Assessment of the Galleries-to-Calories Geobattery

The Galleries-to-Calories Geobattery concept explores the use of abandoned coal mine workings for large-scale thermal energy transport and storage. The system involves injecting waste heat from a supercomputing facility into flooded mine galleries, where groundwater flow can store and transport thermal energy for potential recovery in downgradient district heating and cooling applications. To evaluate the feasibility and performance of the Geobattery under geological and operational uncertainty, we developed a suite of stochastic thermo-hydrological (TH) simulations using Monte Carlo sampling of key uncertain parameters (e.g., permeability, porosity, thermal conductivity, specific heat capacity) and operating conditions (e.g., injection rate, injection temperature). Results identified injection rate and temperature as the most influential parameters governing thermal front propagation, while the geometry of the room-and-pillar structure played a critical role in directing the extent and orientation of thermal advancement. Optimal combinations of material properties for maximizing heat recovery were also determined. To address the high computational cost of coupled-process stochastic modeling, we trained a neural network surrogate model on 24,000 physics-based realizations, achieving an R² > 0.99 and MAE < 0.1 for temperature predictions at monitoring locations. This surrogate enabled an additional 100,000 realizations for global sensitivity analysis and probabilistic thermal resource assessment. The integrated stochastic physics–surrogate modeling framework offers a computationally efficient tool for quantifying uncertainty, identifying key drivers, and informing early-stage design decisions for Geobattery systems.

15 - GEOTHERMAL ENERGY↗

Adaptive Sampling of Time Series During Remote Exploration

This work deals with the challenge of online adaptive data collection in a time series. A remote sensor or explorer agent adapts its rate of data collection in order to track anomalous events while obeying constraints on time and power. This problem is challenging because the agent has limited visibility (all its datapoints lie in the past) and limited control (it can only decide when to collect its next datapoint). This problem is treated from an information-theoretic perspective, fitting a probabilistic model to collected data and optimizing the future sampling strategy to maximize information gain. The performance characteristics of stationary and nonstationary Gaussian process models are compared. Self-throttling sensors could benefit environmental sensor networks and monitoring as well as robotic exploration. Explorer agents can improve performance by adjusting their data collection rate, preserving scarce power or bandwidth resources during uninteresting times while fully covering anomalous events of interest. For example, a remote earthquake sensor could conserve power by limiting its measurements during normal conditions and increasing its cadence during rare earthquake events. A similar capability could improve sensor platforms traversing a fixed trajectory, such as an exploration rover transect or a deep space flyby. These agents can adapt observation times to improve sample coverage during moments of rapid change. An adaptive sampling approach couples sensor autonomy, instrument interpretation, and sampling. The challenge is addressed as an active learning problem, which already has extensive theoretical treatment in the statistics and machine learning literature. A statistical Gaussian process (GP) model is employed to guide sample decisions that maximize information gain. Nonsta tion - ary (e.g., time-varying) covariance relationships permit the system to represent and track local anomalies, in contrast with current GP approaches. Most common GP models are stationary, e.g., the covariance relationships are time-invariant. In such cases, information gain is independent of previously collected data, and the optimal solution can always be computed in advance. Information-optimal sampling of a stationary GP time series thus reduces to even spacing, and such models are not appropriate for tracking localized anomalies. Additionally, GP model inference can be computationally expensive.

Thompson, David R.↗

Digital polycrystalline microstructure generation using diffusion probabilistic models

Accurate micromechanical simulation of polycrystalline materials requires a realistic digital representation of the grain scale microstructure. Here, this work demonstrates the use of a generative diffusion probabilistic model for synthesizing single phase polycrystalline realizations. The model performs well and is capable of producing realistic microstructures consisting of not just simple equiaxed structures but also structures exhibiting more complex spatial arrangements. Masked microstructure generation reveals that the model is context aware of morphological descriptors which may be encoded in the latent space. Training on more diverse data sets, with scaled up architectures, may enable development of future models capable of synthesizing even more complex microstructural features.

36 MATERIALS SCIENCE↗