Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian Evaluation Framework”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Prefire Approach for Probabilistic Assessments of Postfire Debris-Flow Inundation

Increases in wildfire activity and rainfall intensification are driving more postfire debris flows (PFDF) in many regions around the world. PFDFs are most common in the first postfire year and may even occur before a fire is fully controlled. This underscores the importance of assessing postfire hazards before a fire starts. Evaluation of PFDF hazards prior to fire can help strategize interventions lessening the negative effects of future fires. However, debris-flow runout and inundation analyses are not routine in PFDF hazard assessments, partially due to time constraints and substantial uncertainties in boundary conditions. Here, we propose a prefire PFDF inundation assessment framework using a debris-flow runout model based on the Herschel-Bulkley (HB) rheology (HEC-RAS v6.1). We constrain model inputs and parameters using Bayesian posterior analysis, rainfall-runoff simulations, and a debris-flow volume model. We use observations from recent PFDF incidents in northern Arizona, USA, to calibrate model components and then apply our prefire inundation assessment framework in a nearby unburned area. Specifically, we (a) identify yield stress as the most influential factor on inundation extent and arrival time in a HB model, (b) establish posterior distributions for model parameters suitable for forward modeling by leveraging uncertainties in field observations, and (c) implement a predictive forward analysis in an area that has not burned recently to evaluate PFDF inundation under several future fire scenarios. This study improves our ability to assess postfire debris-flow hazards before a fire begins and provides guidance for future applications of single-phase rheological models when assessing PFDF hazards.

54 ENVIRONMENTAL SCIENCES↗

Multi-objective Bayesian alloy design using multi-task Gaussian processes

In design applications, correlations among material properties (such as the tendency for stronger materials to be less ductile) are often neglected. This approach is echoed in multi-objective optimization techniques which treat each performance characteristic as an independent objective, aiming to optimize scalar functions and find optimal Pareto fronts. However, this overlooks the statistical relationships between performance characteristics inherent in a material system. To address this, we propose the use of Bayesian optimization, a highly efficient black-box optimization algorithm known for constructing Gaussian processes (GPs) – uncorrelated surrogates - to model objective functions. Rather than evaluating multiple GPs for each objective function separately, we argue for a shift towards jointly modeling these objective functions, considering their statistical correlations. This integrated approach utilizes naturally occurring relationships among material properties, providing additional information to enhance the performance of the design framework. This requires the replacement of multiple independent GPs with a single multi-task GP, employing a correlation matrix to construct a multi-task kernel function, wherein each task corresponds to a single objective function. Here, we anticipate this refined methodology will better leverage material correlations, improving design optimization results.

36 MATERIALS SCIENCE↗

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING↗

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES↗

Uncertainty quantification in multivariable regression for material property prediction with Bayesian neural networks

With the increased use of data-driven approaches and machine learning-based methods in material science, the importance of reliable uncertainty quantification (UQ) of the predicted variables for informed decision-making cannot be overstated. UQ in material property prediction poses unique challenges, including multi-scale and multi-physics nature of materials, intricate interactions between numerous factors, limited availability of large curated datasets, etc. In this work, we introduce a physics-informed Bayesian Neural Networks (BNNs) approach for UQ, which integrates knowledge from governing laws in materials to guide the models toward physically consistent predictions. To evaluate the approach, we present case studies for predicting the creep rupture life of steel alloys. Experimental validation with three datasets of creep tests demonstrates that this method produces point predictions and uncertainty estimations that are competitive or exceed the performance of conventional UQ methods such as Gaussian Process Regression. Additionally, we evaluate the suitability of employing UQ in an active learning scenario and report competitive performance. The most promising framework for creep life prediction is BNNs based on Markov Chain Monte Carlo approximation of the posterior distribution of network parameters, as it provided more reliable results in comparison to BNNs based on variational inference approximation or related NNs with probabilistic outputs.

36 MATERIALS SCIENCE↗

Dynamic Bayesian Networks for Fault Prognosis

A dynamic Bayesian Network (DBN)-based fault prognosis framework is proposed in this study to predict the future fault probabilities of gradual faults. The proposed framework utilizes the trend in prediction error generated from data driven forecasting models to estimate the future fault beliefs. The accuracy and scalability of the proposed method is evaluated using the data from a Modelica-based virtual testbed. Overall, the developed framework demonstrates good potential in estimating future fault probabilities of gradual faults.

Pradhan, Ojas↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Bayesian Optimization Framework for Imperfect Data or Models

Conventional Bayesian optimization methods implicitly assume that the data and model being optimized are “perfect.” This assumption leads to inaccurate posterior probability distribution functions (PDFs) when applied to “imperfect” data or models. The new Bayesian optimization framework presented in this report provides a way to parameterize the effect of imperfections usually encountered in a prior PDF of generalized data or a model on the posterior PDF. The effects of imperfections are parameterized by a set of constraints imposed on the posterior expectation values of deviations between the data and the model and on their covariance matrix elements. A particular set of values for these constraints conveys an evaluator’s best estimate of the effect of imperfections on the corresponding posterior expectation values. When a prior PDF of generalized data is assumed to be normal, an expression for a posterior PDF satisfying an arbitrary set of constraints is derived analytically for linear models. An analogous iterative algorithm is given for nonlinear models. The corresponding posterior PDF should be used to estimate any posterior expectation values in the presence of imperfections parameterized by that set of constraints. A posterior PDF of a conventional Bayesian optimization method is recovered analytically when all evaluator-specified constraints are set to zero (i.e., in the absence of any imperfections). The analytical expressions derived in this report for normal PDFs and linear models were verified numerically by a Metropolis–Hastings Monte Carlo method. The methods presented herein could be applied to any kind of data or models, including differential cross-section data or integral benchmark experiments.

97 MATHEMATICS AND COMPUTING↗

Optimization of Thermal Conductance at Interfaces Using Machine Learning Algorithms

We report optimization of thermal transport across the interface of two different materials is critical to micro-/nanoscale electronic, photonic, and phononic devices. Although several examples of compositional intermixing at the interfaces having a positive effect on interfacial thermal conductance (ITC) have been reported, an optimum arrangement has not yet been determined because of the large number of potential atomic configurations and the significant computational cost of evaluation. On the other hand, computation-driven materials design efforts are rising in popularity and importance. Yet, the scalability and transferability of machine learning models remain as challenges in creating a complete pipeline for the simulation and analysis of large molecular systems. In this work we present a scalable Bayesian optimization framework, which leverages dynamic spawning of jobs through the Message Passing Interface (MPI) to run multiple parallel molecular dynamics simulations within a parent MPI job to optimize heat transfer at the silicon and aluminum (Si/Al) interface. We found a maximum of 50% increase in the ITC when introducing a two-layer intermixed region that consists of a higher percentage of Si. Because of the random nature of the intermixing, the magnitude of increase in the ITC varies. We observed that both homogeneity/heterogeneity of the intermixing and the intrinsic stochastic nature of molecular dynamics simulations account for the variance in ITC.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bayesian modeling of traffic-related air pollutants: A case study of urban transportation and air quality dynamics in Columbia, South Carolina

Traffic emissions significantly impact near-road air quality and public health. This research applies a Bayesian modeling framework to investigate these impacts using high-resolution traffic and air pollutant data from an urban corridor in Columbia, South Carolina. Despite a data collection period truncated by the COVID-19 lockdown, the Bayesian approach successfully identified significant predictors and quantified model uncertainty. Employing Bayesian Model Selection and Averaging enhanced prediction accuracy and evaluated model uncertainty. Findings indicate that higher temperatures and increased moisture levels elevate particulate matter (PM 1.0 , PM 2.5 , PM 10 ) concentrations, while traffic speed significantly affects nitrogen dioxide (NO 2 ) levels. Specifically, higher average traffic speeds (indicative of smoother flow) correspond to lower NO 2 concentrations, suggesting that less congested conditions reduce NO 2 emissions. This study highlights the robustness of Bayesian methods for generating reliable air quality insights even under data-constrained conditions. The findings underscore the importance of traffic flow management (e.g., reducing congestion) for mitigating near-road NO 2 exposure and provide a basis for developing targeted public health strategies.

54 ENVIRONMENTAL SCIENCES↗

Persistent Sampling: Enhancing the Efficiency of Sequential Monte Carlo

Sequential Monte Carlo (SMC) samplers are powerful tools for Bayesian inference but suffer from high computational costs due to their reliance on large particle ensembles for accurate estimates. We introduce persistent sampling (PS), an extension of SMC that systematically retains and reuses particles from all prior iterations to construct a growing, weighted ensemble. By leveraging multiple importance sampling and resampling from a mixture of historical distributions, PS mitigates the need for excessively large particle counts, directly addressing key limitations of SMC such as particle impoverishment and mode collapse. Crucially, PS achieves this without additional likelihood evaluations-weights for persistent particles are computed using cached likelihood values. This framework not only yields more accurate posterior approximations but also produces marginal likelihood estimates with significantly lower variance, enhancing reliability in model comparison. Furthermore, the persistent ensemble enables efficient adaptation of transition kernels by leveraging a larger, decorrelated particle pool. Experiments on high-dimensional Gaussian mixtures, hierarchical models, and non-convex targets demonstrate that PS consistently outperforms standard SMC and related variants, including recycled and waste-free SMC, achieving substantial reductions in mean squared error for posterior expectations and evidence estimates, all at reduced computational cost. PS thus establishes itself as a robust, scalable, and efficient alternative for complex Bayesian inference tasks.

Karamanis, Minas↗

Development of Automated Atom Probe Tomography capability to study the influence of applied voltage and laser power on the final apparent composition of the analyzed specimen

This study presents the development and implementation of an autonomous Bayesian optimization (BO) framework for controlling and optimizing experimental parameters in Atom Probe Tomography (APT). Using commercial silicon needle samples as a benchmark system, we demonstrate that BO can efficiently navigate the complex parameter space of voltage and laser power to achieve target charge state ratios (specifically Si + /(Si + +Si 2+ )) with minimal experimental evaluations. Our implementation integrates Gaussian Process modeling with the CAMECA atom probe control framework, enabling autonomous adjustment of experimental conditions in real-time. Results show that the algorithm successfully converges to target ratios under different scenarios: maintaining a reference ratio, increasing the ratio (favoring Si 1+ ), and decreasing the ratio (favoring Si 2+ ). The system adapts to specimen evolution during analysis, compensating for changes in apex geometry while maintaining optimization targets. This work establishes a proof of concept for AI-driven optimization in APT, addressing the traditional challenges of manual parameter tuning and paving the way for applications to more complex materials where compositional accuracy is critical.

36 MATERIALS SCIENCE↗

Understanding Model Inadequacy in TRISO Nuclear Fuel Fission Products Release Models: Empirical and Mechanistic Approaches

The increasing use of tristructural isotropic (TRISO) particle fuel in both advanced and existing reactors necessitates a thorough evaluation of uncertainties and shortcomings in TRISO fission product release models. These inadequacies arise from the simplifications made in computational models compared to experimental data. Utilizing the BISON fuel performance code and experimental data from the Advanced Gas Reactor (AGR) program provides a unique chance to rigorously assess these inadequacies within a Bayesian uncertainty quantification (UQ) framework. This study contrasts the standard Bayesian framework with the Kennedy-O'Hagan (KOH) framework, which explicitly accounts for modeling inadequacies, in the context of UQ for TRISO silver release models. It examines both the traditional Arrhenius equation and a more advanced lower-length-scale (LLS)-informed model that incorporates microstructure information. The inverse UQ process applied to AGR-2 and AGR-3/4 datasets identified modeling inadequacy as the primary source of uncertainty, with experimental noise also being significant, while model parameter uncertainty was minimal. Both the Arrhenius and LLS-informed models showed similar levels of modeling inadequacy. For forward predictive UQ using the AGR-1 dataset, the KOH framework enhanced the accuracy and quality of quantified uncertainties by approximately 30% and 40%, respectively, compared to the standard Bayesian framework. This improvement was observed for both the Arrhenius and LLS-informed models. At the engineering scale, both models performed similarly, but the LLS-informed model outperformed the Arrhenius equation at the mesoscale. These findings underscore the importance of explicitly considering modeling inadequacy in the UQ process and highlight the need for ongoing refinement of physics-based models to address these shortcomings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Adaptive Discovery and Mixed-Variable Optimization of Next Generation Synthesizable Microelectronic Materials

Design of new microelectronic materials is characterized by several challenges such as high-dimensionality of the atomic structure-composition variable space, formidable cost of directly using high-fidelity simulations for design optimization, dispersity in literature-reported similar materials and synthesis methods, complex physical mechanisms, and mixed qualitative and quantitative design variables that lead to a disjointed design space. Even though machine learning (ML) techniques have been employed to expedite materials innovation, existing methods treat ML and design optimization as two separate processes, failing to resolve the fundamental challenges associated with high dimensionality and mixed-variable complexity. We have developed a ML enhanced mixed-variable material design optimization framework to efficiently extract useful information from existing data in literature and physics-based simulations to guide the autonomous search for optimal materials. Our proposed framework is composed of four computational modules: (1) a natural language processing (NLP) based virtual screening module, (2) classification based concept exploration module, (3) a density functional theory (DFT)-based high-fidelity evaluation model, and (4) a novel latent-variable Gaussian process (LVGP) ML model for mixed-variable problems with uncertainty quantification, which seamlessly integrates with Bayesian Optimization (BO) and achieves superb efficiency through embedded physics-based dimension reduction. Our approach is demonstrated and validated using the testbed of functional materials exhibiting metal-insulation transitions (MITs), with the targeted reversible resistivity changes (∼10^5) near room temperature. At the end of the 30-month project, we have developed a series of new ML techniques using NLP, conditional variational autoencoders, active learning, latent-variable Gaussian processes, integrated with Bayesian optimization. Our project has resulted in new predicted MITs compounds and improved understanding of MITs microscopic mechanisms, which in turn will revolutionize microelectronics science to provide energy-saving solutions. Our research has improved both creativity and efficiency in transforming rare-event discoveries of new functional materials to persistent innovations. In addition to open-sourcing the online MIT database and the classification model, the LVGP open source code has been downloaded more than 15,000 times within two years. More than 40 MIT compounds have been identified and many have been pursued experimentally via collaborators. The research results are published in close to 20 collaborative papers in high-impact journals, such as Chem. Mater., Appl. Phys. Rev., Sci. Rep., among others of design space.

36 MATERIALS SCIENCE↗

Active operator learning with predictive uncertainty quantification for partial differential equations

With the increased prevalence of neural operators being used to provide rapid solutions to partial differential equations (PDEs), understanding the accuracy of model predictions and the associated error levels is necessary for deploying reliable surrogate models in scientific applications. Existing uncertainty quantification (UQ) frameworks employ ensembles or Bayesian methods, which can incur substantial computational costs during both training and inference. Here, we propose a lightweight predictive UQ method tailored for Deep operator networks (DeepONets) that also generalizes to other operator networks. Numerical experiments on linear and nonlinear PDEs demonstrate that the framework’s uncertainty estimates are unbiased and provide accurate out-of-distribution uncertainty predictions with a sufficiently large training dataset. Our framework provides fast inference and uncertainty estimates that can efficiently drive outer-loop analyses that would be prohibitively expensive with conventional solvers. We demonstrate how predictive uncertainties can be used in the context of Bayesian optimization and active learning problems to yield improvements in accuracy and data-efficiency for outer-loop optimization procedures. In the active learning setup, we extend the framework to Fourier Neural Operators (FNO) and describe a generalized method for other operator networks. To enable real-time deployment, we introduce an inference strategy based on precomputed trunk outputs and a sparse placement matrix, reducing evaluation time by more than a factor of five. Our method provides a practical route to uncertainty-aware operator learning in time-sensitive settings.

97 MATHEMATICS AND COMPUTING↗

Bayes goes fast: Uncertainty quantification for a covariant energy density functional emulated by the reduced basis method

A covariant energy density functional is calibrated using a principled Bayesian statistical framework informed by experimental binding energies and charge radii of several magic and semi-magic nuclei. The Bayesian sampling required for the calibration is enabled by the emulation of the high-fidelity model through the implementation of a reduced basis method (RBM)—a set of dimensionality reduction techniques that can speed up demanding calculations involving partial differential equations by several orders of magnitude. The RBM emulator we build—using only 100 evaluations of the high-fidelity model—is able to accurately reproduce the model calculations in tens of milliseconds on a personal computer, an increase in speed of nearly a factor of 3,300 when compared to the original solver. Besides the analysis of the posterior distribution of parameters, we present model calculations for masses and radii with properly estimated uncertainties. We also analyze the model correlation between the slope of the symmetry energy L and the neutron skin of 48 Ca and 208 Pb. The straightforward implementation and outstanding performance of the RBM makes it an ideal tool for assisting the nuclear theory community in providing reliable estimates with properly quantified uncertainties of physical observables. Such uncertainty quantification tools will become essential given the expected abundance of data from the recently inaugurated and future experimental and observational facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Localization of infrasonic sources via Bayesian back projection

SUMMARY A Bayesian framework is investigated for event-specific localization of infrasonic sources using back projection ray tracing. Direction-of-arrival information from array-based detection analysis is used to initialize a back projection ray path originating from the detecting array location and quantifying propagation characteristics from hypothetical source locations. The Fisher statistic, computed from the array’s beam coherence, is mapped into uncertainty in the launch angles of the ray path. Auxiliary parameters previously introduced for solving the Transport equation to compute geometric spreading along ray paths are used to map uncertainty in the ray launch angles into spatial and temporal uncertainties in the ray path. An atmospheric ensemble approach is applied to account for atmospheric uncertainty, and the relation between uncertainties in the atmospheric state and confidence in estimated localization are evaluated using several ensembles with specified variances. The method is evaluated using a synthetic event in the western United States constructed via forward propagation simulations as well as a single-station, multi-arrival detection from a surface explosion in the western United States. Localization results using this event-specific approach are more accurate and exhibit improved precision than existing Bayesian localization methods that leverage generalized, pre-computed propagation statistics.

58 GEOSCIENCES↗