Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “probabilistic machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Hypothesis Learning in Automated Experiment: Application to Combinatorial Materials Libraries

Machine learning is rapidly becoming an integral part of experimental physical discovery via automated and high-throughput synthesis, and active experiments in scattering and electron/probe microscopy. This, in turn, necessitates the development of active learning methods capable of exploring relevant parameter spaces with the smallest number of steps. In this work, an active learning approach based on conavigation of the hypothesis and experimental spaces is introduced. This is realized by combining the structured Gaussian processes containing probabilistic models of the possible system's behaviors (hypotheses) with reinforcement learning policy refinement (discovery). This approach closely resembles classical human-driven physical discovery, when several alternative hypotheses realized via models with adjustable parameters are tested during an experiment. This approach is demonstrated for exploring concentration-induced phase transitions in combinatorial libraries of Sm-doped BiFeO 3 using piezoresponse force microscopy, but it is straightforward to extend it to higher-dimensional parameter spaces and more complex physical problems once the experimental workflow and hypothesis generation are available.

36 MATERIALS SCIENCE↗

Estimating building occupancy: a machine learning system for day, night, and episodic events

Building occupancy research increasingly emphasizes understanding the social and physical dynamics of how people occupy space. Opportunities in the open source domain including social media, Volunteered Geographic Information, crowdsourcing, and sensor data have proliferated, resulting in the exploration of building occupancy dynamics at varying spatiotemporal scales. At Oak Ridge National Laboratory, research into building occupancies through the development of a global learning framework that accommodates exploitation of open source authoritative sources, including governmental census and surveys, journal articles, real estate databases, and more, to report national and subnational building occupancies across the world continues through the Population Density Tables (PDT) project. This probabilistic learning system accommodates expert knowledge, experience, and open-source data to capture local, socioeconomic, and cultural information about human activity. It does so through a systematic process of data harmonization techniques in the development of observation models for over 50 building types to dynamically update baseline estimates and report probabilistic diurnal and episodic building occupancy estimates. This discussion will explore how PDT is implemented at scale and expanded based on the development of observation model classes and will explain how to interpret and spatially apply the reported probability occupancy estimates and uncertainty.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration

This paper proposes novel noise-free Bayesian optimization strategies that rely on a random exploration step to enhance the accuracy of Gaussian process surrogate models. The new algorithms retain the ease of implementation of the classical GP-UCB algorithm, but the additional random exploration step accelerates their convergence, nearly achieving the optimal convergence rate. Furthermore, to facilitate Bayesian inference with intractable likelihoods, we propose to utilize optimization iterates for maximum a posteriori estimation to build a Gaussian process surrogate model for the unnormalized log-posterior density. We provide bounds for the Hellinger distance between the true and the approximate posterior distributions in terms of the number of design points. We demonstrate the effectiveness of our Bayesian optimization algorithms in nonconvex benchmark objective functions, in a machine learning hyperparameter tuning problem, and in a black-box engineering design problem. The effectiveness of our posterior approximation approach is demonstrated in two Bayesian inference problems for parameters of dynamical systems.

Bayesian inference↗

Stochastic Learning Approach for Binary Optimization: Application to Bayesian Optimal Design of Experiments

Here, we present a novel stochastic approach to binary optimization suited for optimal experimental design (OED) for Bayesian inverse problems governed by mathematical models such as partial differential equations. The OED utility function, namely, the regularized optimality criterion, is cast into a stochastic objective function in the form of an expectation over a multivariate Bernoulli distribution. The probabilistic objective is then solved by using a stochastic optimization routine to find an optimal observational policy. This formulation (a) is generally applicable to binary optimization problems with soft constraints and is ideal for OED and sensor placement problems; (b) does not require differentiability of the original objective function (e.g., a utility function in OED applications) with respect to the design variable, and thus it enables direct employment of sparsity-enforcing penalty functions such as $\ell_0$, without needing to utilize a continuation procedure or apply a rounding technique; (c) exhibits much lower computational cost than traditional gradient-based relaxation approaches; and (d) can be applied to both linear and nonlinear OED problems with proper choice of the utility function. The proposed approach is analyzed from an optimization perspective with detailed convergence analysis of the optimization approach and is also analyzed from a machine learning perspective with correspondence to policy gradient reinforcement learning. The approach is demonstrated numerically by using an idealized two-dimensional Bayesian linear inverse problem and validated by extensive numerical experiments carried out for sensor placement in a parameter identification setup.

97 MATHEMATICS AND COMPUTING↗

COVID-19 dynamics across the US: A deep learning study of human mobility and social behavior

This paper presents a deep learning framework for epidemiology system identification from noisy and sparse observations with quantified uncertainty. The proposed approach employs an ensemble of deep neural networks to infer the time-dependent reproduction number of an infectious disease by formulating a tensor-based multi-step loss function that allows us to efficiently calibrate the model on multiple observed trajectories. The method is applied to a mobility and social behavior-based SEIR model of COVID-19 spread. The model is trained on Google and Unacast mobility data spanning a period of 66 days, and is able to yield accurate future forecasts of COVID-19 spread in 203 US counties within a time-window of 15 days. Interestingly, a sensitivity analysis that assesses the importance of different mobility and social behavior parameters reveals that attendance of close places, including workplaces, residential, and retail and recreational locations, has the largest impact on the effective reproduction number. Furthermore, the model enables us to rapidly probe and quantify the effects of government interventions, such as lock-down and re-opening strategies. Taken together, the proposed framework provides a robust workflow for data-driven epidemiology model discovery under uncertainty and produces probabilistic forecasts for the evolution of a pandemic that can judiciously provide information for policy and decision making. All codes and data accompanying this manuscript are available at https://github.com/PredictiveIntelligenceLab/DeepCOVID19.

60 APPLIED LIFE SCIENCES↗

Emulation With Uncertainty Quantification of Regional Sea‐Level Change Caused by the Antarctic Ice Sheet

Abstract Projecting regional sea‐level change under various climate‐change scenarios typically involves running forward simulations of the Earth's gravitational, rotational and deformational (GRD) response to ice‐mass change, which requires substantial computational cost if applied to probabilistic frameworks requiring thousands to millions of samples. Here we build emulators of regional sea‐level change at 27 coastal locations, due to the GRD effects associated with future Antarctic Ice Sheet mass change over the 21st century. The emulators are evaluated against a numerical sea‐level model applied to an ensemble of ice‐sheet model simulations of the Antarctic Ice Sheet through 2100. We build a physics‐based emulator using a recent sensitivity kernel approach and compare it to machine learning based emulators (neural network and conditional variational autoencoder methods). In order to quantify uncertainty, we derive well‐calibrated prediction intervals for regional sea‐level change via split‐conformal inference and linear regression, and show that Monte Carlo dropout does not yield well‐calibrated uncertainties in this instance. We also demonstrate substantial gains in computational efficiency using both the physics‐based emulator and neural networks in comparison to the numerical model for the complete regional sea‐level solution. Overall, we find the physics‐based emulator modestly outperforms the machine learning emulators for this problem.

58 GEOSCIENCES↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Insights into Cation Ordering of Double Perovskite Oxides from Machine Learning and Causal Relations

This work investigates origins of cation ordering in double perovskites using first-principles theory computations combined with machine learning (ML) and causal relations. We have considered various oxidation states of A, A', B, and B' from the family of transition metal ions to construct a diverse compositional space. A conventional framework employing traditional ML classification algorithms such as Random Forest (RF) coupled with appropriate features including geometry-driven and key structural modes leads to accurate prediction (~98%) of A-site cation ordering. We have evaluated the accuracy of ML models by employing analyses of decision paths, assignments of probabilistic confidence bound, and finally a direct non-Gaussian acyclic structural equation model to investigate causality. Our study suggests that structural modes are crucial for classifying layered, columnar, and rock-salt ordering. The charge difference between A and A' is the most important feature for predicting clear layered ordering, which in turn depends on the B and B' charge separation. We have also designed mathematical relationships with these features to derive energy differences to form clear layered ordering. Here, the trilinear coupling between tilt, in-phase rotation, and A-site antiferroelectric displacement in the Landau free-energy expansion becomes the necessary condition behind formation of A-site cation ordering.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ClimGen: Learning the Forcing-Response Relationship in Climate System

Solar Radiation Management (SRM) is emerging as a potential geoengineering strategy to address the anthropogenic impact on climate, but its effective implementation requires an iterative and large ensemble of highly accurate and efficient climate projections. Traditional climate projections rely on executing computationally demanding and time-consuming numerical climate models. Recent advances in machine learning (ML) aim to enhance these approaches by emulating traditional methods. In this work, we propose a novel framework for directly learning the relationship between solar radiation flux at the top of the atmosphere and the corresponding surface temperature response. To evaluate the feasibility of this direct ML-based projection, we developed a dataset using an intermediate complexity model, incorporating a comprehensive suite of different forcing patterns and evaluation metrics to rigorously assess the ML model’s performance. We introduce a Conditional Denoising Diffusion Probabilistic Model (cDDPM) for this task, which demonstrates encouraging skill in representing climate statistics under previously unseen forcing patterns. This approach provides a promising pathway for direct climate projections by accurately learning the forcing-response relationship, with a wide range of applications in impact mitigation, emissions policy design, and SRM strategies.

Chen, Tse-Chun [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

A Probabilistic Reasoner Based on Bayes Risk for Damage Detection in Structural Systems

Structural health monitoring (SHM) systems are used to inform operation of structural systems subject to loads and environments that may affect their integrity. SHM systems rely on continuous monitoring of the structure to determine its health state. These systems are often coupled with a model of the deployed structure to determine the consequences of changes in the system by forecasting the response to future states. These models, which may be thought of as digital twins, need to be updated to reflect the latest state of the structural system. This work makes use of an uncertainty-aware machine learning model that enforces distance preservation of the original input space to determine deviations from the training data input space distributions. This workflow enables domain shift detection to determine whether damage is present in the structure. The uncertainty metrics generated by this network are then used in a Bayes risk framework to design an optimal damage detector given cost and risk considerations. The approach is demonstrated on a computational example with simulated damage.

Najera-Flores, David [ATA Engineering, Inc.]↗

Optimal Transport as a Tool for Scientific Discovery in Radiation Biology

This report summarizes findings from research conducted for the “Exploration of the Poten tial for Artificial Intelligence and Machine Learning to Advance Low-Dose Radiation Biology Re search” (RadBio-AI) program, supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, under Awards KP1601011/FWP CC121 and KP1601017/FWP CC121. The research reported here was undertaken in an effort to assess the potential of optimal measure transport methods as components within the larger scope of a com putational framework envisioned to support research in the radiation biology domain. Within this effort, our interest centered on enabling a unified generic framework where probabilistic modeling, inference, and statistical learning can be carried out for a wide range of data distributions. As described next in Section 1 (and in more detail in our original publication), optimal measure transport offers the possibility of such unified approach.

97 MATHEMATICS AND COMPUTING↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Electrocardiographic changes predate Parkinson’s disease onset

Autonomic nervous system involvement precedes the motor features of Parkinson’s disease (PD). Our goal was to develop a proof-of-concept model for identifying subjects at high risk of developing PD by analysis of cardiac electrical activity. We used standard 10-s electrocardiogram (ECG) recordings of 60 subjects from the Honolulu Asia Aging Study including 10 with prevalent PD, 25 with prodromal PD, and 25 controls who never developed PD. Various methods were implemented to extract features from ECGs including simple heart rate variability (HRV) metrics, commonly used signal processing methods, and a Probabilistic Symbolic Pattern Recognition (PSPR) method. Extracted features were analyzed via stepwise logistic regression to distinguish between prodromal cases and controls. Stepwise logistic regression selected four features from PSPR as predictors of PD. The final regression model built on the entire dataset provided an area under receiver operating characteristics curve (AUC) with 95% confidence interval of 0.90 [0.80, 0.99]. The five-fold cross-validation process produced an average AUC of 0.835 [0.831, 0.839]. We conclude that cardiac electrical activity provides important information about the likelihood of future PD not captured by classical HRV metrics. Machine learning applied to ECGs may help identify subjects at high risk of having prodromal PD.

59 BASIC BIOLOGICAL SCIENCES↗

Learning from learning machines: improving the predictive power of energy-water-land nexus models with insights from complex measured and simulated data

Focal Area(s): Insights gleaned from complex data (observed and simulated) using AI; Predictive modeling through the use of AI techniques and AI-derived model components, including physics- and knowledge-informed models; Energy-water-land nexus and integrated energy systems – models of MultiSector dynamics. Science Challenge: Scientific communities are in need of tools for the computational integration of physics-based models, experimental data, and empirical/observational studies across a broad range of temporal and spatial scales to explore and meet EESSD grand challenges. We require AI algorithms for the discovery of process-drivers in the Earth-energy-human system and to link them with complex measured and simulated data. We aim to understand the energy-water-land nexus, particularly under extreme forcing scenarios and rare events, leveraging scale-aware AI process models, including probabilistic uncertainties, and benefiting from the emerging 5G-enabled landscape-scale sensing and edge computing capabilities. These models are needed to identify instabilities and tipping points that manifest extreme system behaviors with consequences for integrated energy systems and the environment. This work requires fundamental advances in uncertainty quantification, in particular, to identify and model unlikely but catastrophic outliers. Specifically, we need models that are interpretable to domain scientists and, essentially, explainable to public and private stake holders.

54 ENVIRONMENTAL SCIENCES↗

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗