Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Quantum Chemistry-Informed Active Learning to Accelerate the Design and Discovery of Sustainable Energy Storage Materials

Here we employed Density Functional Theory (DFT) to compute oxidation potentials of 1,400 homobenzylic ether molecules to search for the ideal sustainable redoxmer design. The generated data were used to construct an active learning model based on Bayesian optimization (BO) that targets candidates with desired oxidation potentials utilizing only a minimal number of DFT calculations. The active learning model demonstrated not only significant efficiency improvement over the random selection approach but also robust capability in identifying desired candidates in an untested set of 112,000 homobenzylic ether molecules. Our findings highlight the efficacy of quantum chemistry-informed active learning to accelerate the discovery of materials with desired properties from a vast chemical space.

25 ENERGY STORAGE↗

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark -- A Bayesian Inverse UQ-based Approach for Data Assimilation

The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. When comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that GLLS predictions fail to replicate computed response distributions for nonlinear applications, while MOCABA shows near agreement, and IUQ uses computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

FOS: Computer and information sciences↗

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark—A Bayesian Inverse UQ-Based Approach for Data Assimilation

The Organisation for Economic Co-operation and Development Working Party on Nuclear Criticality Safety has proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian inverse uncertainty quantification (IUQ) employing scientific machine learning surrogate models as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of generalized linear least squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. Here, when comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that the GLLS predictions failed to replicate the computed response distributions for nonlinear applications, while MOCABA showed near agreement, and IUQ used the computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

Bayesian calibration↗

Exploring the contamination of the DES-Y1 cluster sample with SPT-SZ selected clusters

ABSTRACT We perform a cross validation of the cluster catalogue selected by the red-sequence Matched-filter Probabilistic Percolation algorithm (redMaPPer) in Dark Energy Survey year 1 (DES-Y1) data by matching it with the Sunyaev–Zel’dovich effect (SZE) selected cluster catalogue from the South Pole Telescope SPT-SZ survey. Of the 1005 redMaPPer selected clusters with measured richness $\hat{\lambda }\gt 40$ in the joint footprint, 207 are confirmed by SPT-SZ. Using the mass information from the SZE signal, we calibrate the richness–mass relation using a Bayesian cluster population model. We find a mass trend λ ∝ MB consistent with a linear relation (B ∼ 1), no significant redshift evolution and an intrinsic scatter in richness of σλ = 0.22 ± 0.06. By considering two error models, we explore the impact of projection effects on the richness–mass modelling, confirming that such effects are not detectable at the current level of systematic uncertainties. At low richness SPT-SZ confirms fewer redMaPPer clusters than expected. We interpret this richness dependent deficit in confirmed systems as due to the increased presence at low richness of low-mass objects not correctly accounted for by our richness-mass scatter model, which we call contaminants. At a richness $\hat{\lambda }=40$, this population makes up ${\gt}12{{\ \rm per\ cent}}$ (97.5 percentile) of the total population. Extrapolating this to a measured richness $\hat{\lambda }=20$ yields ${\gt}22{{\ \rm per\ cent}}$ (97.5 percentile). With these contamination fractions, the predicted redMaPPer number counts in different plausible cosmologies are compatible with the measured abundance. The presence of such a population is also a plausible explanation for the different mass trends (B ∼ 0.75) obtained from mass calibration using purely optically selected clusters. The mean mass from stacked weak lensing (WL) measurements suggests that these low-mass contaminants are galaxy groups with masses ∼3–5 × 1013 M⊙ which are beyond the sensitivity of current SZE and X-ray surveys but a natural target for SPT-3G and eROSITA.

79 ASTRONOMY AND ASTROPHYSICS↗

Bayesian model averaging for analysis of lattice field theory results

Statistical modeling is a key component in the extraction of physical results from lattice field theory calculations. Although the general models used are often strongly motivated by physics, many model variations can frequently be considered for the same lattice data. Model averaging, which amounts to a probability-weighted average over all model variations, can incorporate systematic errors associated with model choice without being overly conservative. We discuss the framework of model averaging from the perspective of Bayesian statistics, and give useful formulae and approximations for the particular case of least-squares fitting, commonly used in modeling lattice results. In addition, we frame the common problem of data subset selection (e.g. choice of minimum and maximum time separation for fitting a two-point correlation function) as a model selection problem and study model averaging as a straightforward alternative to manual selection of fit ranges. Numerical examples involving both mock and real lattice data are given.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

CODEX weak lensing mass catalogue and implications on the mass–richness relation

The COnstrain Dark Energy with X-ray clusters (CODEX) sample contains the largest flux limited sample of X-ray clusters at 0.35 < z < 0.65. It was selected from ROSAT data in the 10 000 square degrees of overlap with BOSS, mapping a total number of 2770 high-z galaxy clusters. We present here the full results of the CFHT CODEX programme on cluster mass measurement, including a reanalysis of CFHTLS Wide data, with 25 individual lensing-constrained cluster masses. We employ lensfit shape measurement and perform a conservative colour–space selection and weighting of background galaxies. Using the combination of shape noise and an analytic covariance for intrinsic variations of cluster profiles at fixed mass due to large-scale structure, miscentring, and variations in concentration and ellipticity, we determine the likelihood of the observed shear signal as a function of true mass for each cluster. We combine 25 individual cluster mass likelihoods in a Bayesian hierarchical scheme with the inclusion of optical and X-ray selection functions to derive constraints on the slope α, normalization β, and scatter σln λ|μ of our richness–mass scaling relation model in log-space: ${\langle {\rm In}\,\, \lambda\!\!\mid\!\!\mu\rangle = \alpha\mu + \beta,} $ with μ = ln (M 200c /M piv ), and M piv = 10 14.81 M ⊙ . We find a slope $\alpha = 0.49^{+0.20}_{-0.15}$, normalization $\exp (\beta) = 84.0^{+9.2}_{-14.8}$, and $\sigma _{\ln \lambda | \mu } = 0.17^{+0.13}_{-0.09}$ using CFHT richness estimates. In comparison to other weak lensing richness–mass relations, we find the normalization of the richness statistically agreeing with the normalization of other scaling relations from a broad redshift range (0.0 < z < 0.65) and with different cluster selection (X-ray, Sunyaev–Zeldovich, and optical).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

BOOTS: Bayesian Optimization for Optimal Test Selection

BOOTS is a package for optimal selection of candidate operating points in large-scale manufacturing applications. It is designed to optimally select operating points to maximize predicted values of product quality and resource efficiency according to a data-driven model. The package is based on the use of a multi-input multi-output (MIMO) Gaussian process model to describe the relationships between inputs (operating points) and outputs (product quality and resource efficiency). This package does not provide any site- or process-specific information.

Villez, Kris [Oak Ridge National Laboratory (ORNL)↗

Hierarchical Bayesian modeling for Inverse Uncertainty Quantification of system thermal-hydraulics code using critical flow experimental data

The best estimate plus uncertainty methodology in nuclear system thermal-hydraulic studies necessitates a comprehensive understanding of uncertainties in system code predictions. The forward uncertainty quantification (UQ) process involves the propagation of input uncertainties through the computational models to obtain uncertainties in the outputs. To this end, achieving an accurate estimation of input uncertainties is important, which is the focus of inverse UQ (IUQ). Traditionally, research in Bayesian IUQ within the nuclear engineering domain has largely relied on single-level Bayesian inference. While being effective for relatively small datasets, this approach encounters limitations for cases with large datasets. The use of a single-level model may prove inefficient, as the resultant posterior distributions can significantly differ when distinct subsets of data are employed. To address this issue, we employ an hierarchical Bayesian model for IUQ. Furthermore, this approach involves organizing observations into different groups based on the test conditions, thereby accommodating varying calibration parameters across these distinct groups. In this study, we developed and implemented a hierarchical Bayesian IUQ method to consider the grouping effect of critical flow measurement data from various geometries. Comparing the outcomes of IUQ under different selections of test data using hierarchical Bayesian IUQ against those obtained from single-level Bayesian IUQ, the forward propagation of hierarchical Bayesian IUQ results demonstrates a notably improved agreement with the experimental data.

42 ENGINEERING↗

Bayesian stability and force modeling for uncertain machining processes

Accurately simulating machining operations requires knowledge of the cutting force model and system frequency response. However, this data is collected using specialized instruments in an ex-situ manner. Bayesian statistical methods instead learn the system parameters using cutting test data, but to date, these approaches have only considered milling stability. This paper presents a physics-based Bayesian framework which incorporates both spindle power and milling stability. Initial probabilistic descriptions of the system parameters are propagated through a set of physics functions to form probabilistic predictions about the milling process. The system parameters are then updated using automatically selected cutting tests to reduce parameter uncertainty and identify more productive cutting conditions, where spindle power measurements are used to learn the cutting force model. The framework is demonstrated through both numerical and experimental case studies. Results show that the approach accurately identifies both the system natural frequency and cutting force model.

42 ENGINEERING↗

Sequential Bayesian Experimental Design for Calibration of Expensive Simulation Models

Simulation models of critical systems often have parameters that need to be calibrated using observed data. For expensive simulation models, calibration is done using an emulator of the simulation model built on simulation output at different parameter settings. Using intelligent and adaptive selection of parameters to build the emulator can drastically improve the efficiency of the calibration process. The article proposes a sequential framework with a novel criterion for parameter selection that targets learning the posterior density of the parameters. The emergent behavior from this criterion is that exploration happens by selecting parameters in uncertain posterior regions while simultaneously exploitation happens by selecting parameters in regions of high posterior density. Furthermore, the advantages of the proposed method are illustrated using several simulation experiments and a nuclear physics reaction model.

97 MATHEMATICS AND COMPUTING↗

AL4GAP: Active learning workflow for generating DFT-SCAN accurate machine-learning potentials for combinatorial molten salt mixtures

Machine learning interatomic potentials have emerged as a powerful tool for bypassing the spatiotemporal limitations of ab initio simulations, but major challenges remain in their efficient parameterization. We present AL4GAP, an ensemble active learning software workflow for generating multicomposition Gaussian approximation potentials (GAP) for arbitrary molten salt mixtures. The workflow capabilities include: (1) setting up user-defined combinatorial chemical spaces of charge neutral mixtures of arbitrary molten mixtures spanning 11 cations (Li, Na, K, Rb, Cs, Mg, Ca, Sr, Ba and two heavy species, Nd, and Th) and 4 anions (F, Cl, Br, and I), (2) configurational sampling using low-cost empirical parameterizations, (3) active learning for down-selecting configurational samples for single point density functional theory calculations at the level of Strongly Constrained and Appropriately Normed (SCAN) exchange-correlation functional, and (4) Bayesian optimization for hyperparameter tuning of two-body and many-body GAP models. Here, we apply the AL4GAP workflow to showcase high throughput generation of five independent GAP models for multicomposition binary-mixture melts, each of increasing complexity with respect to charge valency and electronic structure, namely: LiCl–KCl, NaCl–CaCl 2 , KCl–NdCl 3 , CaCl 2 –NdCl 3 , and KCl–ThCl 4 . Our results indicate that GAP models can accurately predict structure for diverse molten salt mixture with density functional theory (DFT)-SCAN accuracy, capturing the intermediate range ordering characteristic of the multivalent cationic melts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Search for very-short-baseline oscillations of reactor antineutrinos with the SoLid detector

In this paper we report the first scientific result based on antineutrinos emitted from the BR2 reactor at SCK CEN. The SoLid experiment uses a novel type of highly granular detector whose basic detection unit combines two scintillators, polyvinyl toluene (PVT) and Li 6 F : ZnS ( Ag ) , to measure antineutrinos via their inverse- β -decay products. An advantage of PVT is its highly linear response as a function of deposited particle energy. The full-scale detector comprises 12 800 voxels and operates over a very short 6.3–8.9 m baseline from the reactor core. The detector segmentation and its three-dimensional imaging capabilities facilitate the extraction of the positron energy from the rest of the visible energy, allowing the latter to be utilized for signal-background discrimination. We present a result obtained from 280 reactor-on days (55 MW mean power) and 172 reactor-off days, respectively, of live data taking. A total of 29 479 ± 603 (stat) antineutrino candidates have been selected, corresponding to an average rate of 105 events per day and a signal-to-background ratio of 0.27. A search for disappearance of antineutrinos to a sterile state has been conducted using complementary model-dependent frequentist and Bayesian fits, providing constraints on the allowed region of the reactor antineutrino anomaly.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sparsifying priors for Bayesian uncertainty quantification in model discovery

We propose a probabilistic model discovery method for identifying ordinary differential equations governing the dynamics of observed multivariate data. Our method is based on the sparse identification of nonlinear dynamics (SINDy) framework, where models are expressed as sparse linear combinations of pre-specified candidate functions. Promoting parsimony through sparsity leads to interpretable models that generalize to unknown data. Instead of targeting point estimates of the SINDy coefficients, we estimate these coefficients via sparse Bayesian inference. The resulting method, uncertainty quantification SINDy (UQ-SINDy), quantifies not only the uncertainty in the values of the SINDy coefficients due to observation errors and limited data, but also the probability of inclusion of each candidate function in the linear combination. UQ-SINDy promotes robustness against observation noise and limited data, interpretability (in terms of model selection and inclusion probabilities) and generalization capacity for out-of-sample forecast. Sparse inference for UQ-SINDy employs Markov chain Monte Carlo, and we explore two sparsifying priors: the spike and slab prior, and the regularized horseshoe prior. UQ-SINDy is shown to discover accurate models in the presence of noise and with orders-of-magnitude less data than current model discovery methods, thus providing a transformative method for real-world applications which have limited data.

97 MATHEMATICS AND COMPUTING↗

Weak Lensing Mass Calibration of the ACT DR5 Galaxy Clusters with the DES Year 3 Weak Lensing Data

We use weak gravitational lensing measurements from Year 3 Dark Energy Survey data to calibrate the masses of 443 galaxy clusters selected via the Sunyaev-Zel'dovich effect from Atacama Cosmology Telescope Data Release 5 maps of the cosmic microwave background. We incorporate redshift and SZ measurements for individual clusters into a hierarchical model for the stacked lensing signals and perform Bayesian analyses to constrain the hydrostatic mass bias of the clusters. Our treatment of systematic uncertainties includes a prescription for measuring and accounting for the weak lensing boost factor, consideration of a miscentering effect, as well as marginalization over uncertainties in the source galaxy photometric redshift distributions and shear calibration. The resultant constraints on the normalization of the mass-observable relation have a precision of approximately 7%, with the mean WL halo mass of M $_{500c}$ = 5.4 × 10$^{14}$ M $_{⊙}$. We measure the bias between the true cluster mass and the mass estimated from the SZ signal based on an X-ray-calibrated scaling relation assuming hydrostatic equilibrium, to be 1 - b = 0.74$^{+0.06}$ $_{-0.05}$ over the full sample. When splitting the clusters into high (z = 0.43-0.70) and low (z = 0.15-0.43) redshift bins, we measure 1 - b = 0.58$^{+0.06}$ $_{-0.05}$ and 0.81$^{+0.08}$ $_{-0.06}$, respectively. When introducing additional freedom in redshift and mass to the hydrostatic bias model, we find that 1 - b decreases with redshift (with the power law of -1.8$^{+0.5}$ $_{-0.6}$, 99.95% confidence), consistent with findings from other recent studies, while we do not find any significant trend in mass. We also demonstrate that our result is robust against various systematics such as a scale cut, priors on baryonic and miscentering parameters, and degree of scatter in mass-observable relation. The weak-lensing mass calibration presented in this study will be a useful tool for using the ACT clusters as probes of astrophysics, and as a step towards using their abundance as a cosmological probe.

Shin, T. [Carnegie Mellon U.] (ORCID:0000000263895↗

Active Galactic Nuclei Continuum Reverberation Mapping Based on Zwicky Transient Facility Light Curves

We perform a systematic survey of active galactic nuclei (AGNs) continuum lags using ~3 days cadence gri-band light curves from the Zwicky Transient Facility. We select a sample of 94 type 1 AGNs at z < 0.8 with significant and consistent inter-band lags based on the interpolated cross-correlation function method and the Bayesian method JAVELIN. Within the framework of the "lamp-post" reprocessing model, our findings are: (1) The continuum emission (CE) sizes inferred from the data are larger than the disk sizes predicted by the standard thin-disk model. (2) For a subset of the sample, the CE size exceeds the theoretical limit of the self-gravity radius (12 lt-days) for geometrically thin disks. (3) The CE size scales with continuum luminosity as R CE ∝ L 0.48±0.04 with a scatter of 0.2 dex, analogous to the well-known radius–luminosity relation of broad Hβ. These findings suggest a significant contribution of diffuse continuum emission from the broad-line region (BLR) to AGN continuum lags. We find that the R CE –L relation can be explained by a photoionization model that assumes ~23% of the total flux comes from the diffuse BLR emission. In addition, the ratio of the CE size and model-predicted disk size anticorrelates with the continuum luminosity, which is indicative of a potential nondisk BLR lag contribution evolving with the luminosity. Finally, a robust positive correlation between the CE size and black hole mass is detected.

79 ASTRONOMY AND ASTROPHYSICS↗

Automating Bayesian inference and design to quantify acoustic particle levitation

Self-propulsion of micro- and nanoparticles powered by ultrasound provides an attractive strategy for the remote manipulation of colloidal matter using biocompatible energy inputs. Quantitative understanding of particle motion and its dependence on size, shape, and composition requires accurate characterization of the acoustic field, which depends sensitively on the experimental setup. Here, we show how automated experiments based on Bayesian inference and design can accurately and efficiently characterize the acoustic field within resonant chambers used to propel acoustic nanomotors. Repeated cycles of observation, inference, and design (OID) are guided by a physical model that describes the rate at which levitating particles approach the nodal plane. Using video microscopy, we observe the relaxation of tracer particles to this plane following the application of the acoustic field. We use sequential Monte Carlo methods to infer model parameters such as the amplitude and frequency of the resonant chamber while accounting for particle-level measurement noise and population-level heterogeneity in the field. Guided by simulated outcomes, we select the optimal design for the next experiment as to maximize the information gain in the relevant parameters. We show how this iterative process serves to discriminate between competing hypotheses and efficiently converges to accurate parameter estimates using only few automated experiments. We discuss the need for model criticism to ensure the validity of the guiding model throughout automated cycles of observation, inference, and design. Furthermore, this work demonstrates how Bayesian methods can learn the parameters of nonlinear, hierarchical models used to describe video microscopy data of active colloids.

42 ENGINEERING↗

The dark energy survey 5-yr photometrically identified type Ia supernovae

ABSTRACT As part of the cosmology analysis using Type Ia Supernovae (SN Ia) in the Dark Energy Survey (DES), we present photometrically identified SN Ia samples using multiband light curves and host galaxy redshifts. For this analysis, we use the photometric classification framework SuperNNovatrained on realistic DES-like simulations. For reliable classification, we process the DES SN programme (DES-SN) data and introduce improvements to the classifier architecture, obtaining classification accuracies of more than 98 per cent on simulations. This is the first SN classification to make use of ensemble methods, resulting in more robust samples. Using photometry, host galaxy redshifts, and a classification probability requirement, we identify 1863 SNe Ia from which we select 1484 cosmology-grade SNe Ia spanning the redshift range of 0.07 < z < 1.14. We find good agreement between the light-curve properties of the photometrically selected sample and simulations. Additionally, we create similar SN Ia samples using two types of Bayesian Neural Network classifiers that provide uncertainties on the classification probabilities. We test the feasibility of using these uncertainties as indicators for out-of-distribution candidates and model confidence. Finally, we discuss the implications of photometric samples and classification methods for future surveys such as Vera C. Rubin Observatory Legacy Survey of Space and Time.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Forecasting Multi-Wave Epidemics Through Bayesian Inference

We present a simple, near-real-time Bayesian method to infer and forecast a multiwave outbreak, and demonstrate it on the COVID-19 pandemic. The approach uses timely epidemiological data that has been widely available for COVID-19. It provides short-term forecasts of the outbreak’s evolution, which can then be used for medical resource planning. The method postulates one- and multiwave infection models, which are convolved with the incubation-period distribution to yield competing disease models. The disease models’ parameters are estimated via Markov chain Monte Carlo sampling and information-theoretic criteria are used to select between them for use in forecasting. The method is demonstrated on two- and three-wave COVID-19 outbreaks in California, New Mexico and Florida, as observed during Summer-Winter 2020. We find that the method is robust to noise, provides useful forecasts (along with uncertainty bounds) and that it reliably detected when the initial single-wave COVID-19 outbreaks transformed into successive surges as containment efforts in these states failed by the end of Spring 2020.

59 BASIC BIOLOGICAL SCIENCES↗