Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie↗

Automated integration gate selection for Gaussian mixture model pulse shape discrimination

Pulse shapes differ between neutron and gamma particles when measured with detector devices employing pulse shape discriminating (PSD) scintillators. Digitized waveforms can be used in detection systems to perform pulse shape discrimination for this application. Prior Gaussian Mixture Model (GMM) methods require access to the pulse full-waveform. Reducing the waveform to a smaller set of combined samples reduces computational cost while affecting PSD performance. In this work, we develop a method for selecting the best performing combination of integration gates, or contiguous summed segments of the digitized pulse for PSD. The method uses a discrimination score based on the GMM PSD approach. Furthermore, the final selection is performed using Bayesian Optimization. PSD detection results are compared with varying numbers of selected gates on time-of-flight (TOF) data. This method can be used to fully automate the selection of gates in an unsupervised (without ground truth labels) setting.

42 ENGINEERING↗

Synthetic Streamflow Datasets Derived from DOE 9505 for Select Texas Basins

This dataset is generated using a Bayesian Hidden Markov Model trained on the DOE 9505 streamflow projection ensemble. A set of 21,000 streamflow realizations are generated for the Colorado, Sabine, and Trinity river basins in Texas. Separate models are trained either using the full 9505 ensemble or a subset based on three hyperparameters: bias correction, downscaling, and hydrological model.

drought↗

Detection and Modeling of High-Dimensional Thresholds for Fault Detection and Diagnosis

Many Fault Detection and Diagnosis (FDD) systems use discrete models for detection and reasoning. To obtain categorical values like oil pressure too high, analog sensor values need to be discretized using a suitablethreshold. Time series of analog and discrete sensor readings are processed and discretized as they come in. This task isusually performed by the wrapper code'' of the FDD system, together with signal preprocessing and filtering. In practice,selecting the right threshold is very difficult, because it heavily influences the quality of diagnosis. If a threshold causesthe alarm trigger even in nominal situations, false alarms will be the consequence. On the other hand, if threshold settingdoes not trigger in case of an off-nominal condition, important alarms might be missed, potentially causing hazardoussituations. In this paper, we will in detail describe the underlying statistical modeling techniques and algorithm as well as the Bayesian method for selecting the most likely shape and its parameters. Our approach will be illustrated by several examples from the Aerospace domain.

Computer Systems↗

Object-Oriented Bayesian Networks (OOBN) for Aviation Accident Modeling and Technology Portfolio Impact Assessment

The concern for reducing aviation safety risk is rising as the National Airspace System in the United States transforms to the Next Generation Air Transportation System (NextGen). The NASA Aviation Safety Program is committed to developing an effective aviation safety technology portfolio to meet the challenges of this transformation and to mitigate relevant safety risks. The paper focuses on the reasoning of selecting Object-Oriented Bayesian Networks (OOBN) as the technique and commercial software for the accident modeling and portfolio assessment. To illustrate the benefits of OOBN in a large and complex aviation accident model, the in-flight Loss-of-Control Accident Framework (LOCAF) constructed as an influence diagram is presented. An OOBN approach not only simplifies construction and maintenance of complex causal networks for the modelers, but also offers a well-organized hierarchical network that is easier for decision makers to exploit the model examining the effectiveness of risk mitigation strategies through technology insertions.

Shih, Ann T.↗

The SRG/eROSITA All-Sky Survey: Dark Energy Survey year 3 weak gravitational lensing by eRASS1 selected galaxy clusters

Context. Number counts of galaxy clusters across redshift are a powerful cosmological probe if a precise and accurate reconstruction of the underlying mass distribution is performed – a challenge called mass calibration. With the advent of wide and deep photometric surveys, weak gravitational lensing (WL) by clusters has become the method of choice for this measurement. Aims. We measured and validated the WL signature in the shape of galaxies observed in the first three years of the Dark Energy Survey (DES Y3) caused by galaxy clusters and groups selected in the first all-sky survey performed by SRG (Spectrum Roentgen Gamma)/eROSITA (eRASS1). These data were then used to determine the scaling between the X-ray photon count rate of the clusters and their halo mass and redshift. Methods. We empirically determined the degree of cluster member contamination in our background source sample. The individual cluster shear profiles were then analyzed with a Bayesian population model that self-consistently accounts for the lens sample selection and contamination and includes marginalization over a host of instrumental and astrophysical systematics. To quantify the accuracy of the mass extraction of that model, we performed mass measurements on mock cluster catalogs with realistic synthetic shear profiles. This allowed us to establish that hydrodynamical modeling uncertainties at low lens redshifts (z < 0.6) are the dominant systematic limitation. At high lens redshift, the uncertainties of the sources’ photometric redshift calibration dominate. Results. With regard to the X-ray count rate to halo mass relation, we determined its amplitude, its mass trend, the redshift evolution of the mass trend, the deviation from self-similar redshift evolution, and the intrinsic scatter around this relation. Conclusions. The mass calibration analysis performed here sets the stage for a joint analysis with the number counts of eRASS1 clusters to constrain a host of cosmological parameters. We demonstrate that WL mass calibration of galaxy clusters can be performed successfully with source galaxies whose calibration was performed primarily for cosmic shear experiments, opening the way for the cluster cosmological exploitation of future optical and NIR surveys like Euclid and LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

Developing and applying quantifiable metrics for diagnostic and experiment design on Z

This project applies methods in Bayesian inference and modern statistical methods to quantify the value of new experimental data, in the form of new or modified diagnostic configurations and/or experiment designs. We demonstrate experiment design methods that can be used to identify the highest priority diagnostic improvements or experimental data to obtain in order to reduce uncertainties on critical inferred experimental quantities and select the best course of action to distinguish between competing physical models. Bayesian statistics and information theory provide the foundation for developing the necessary metrics, using two high impact experimental platforms on Z as exemplars to develop and illustrate the technique. We emphasize that the general methodology is extensible to new diagnostics (provided synthetic models are available), as well as additional platforms. We also discuss initial scoping of additional applications that began development in the last year of this LDRD.

97 MATHEMATICS AND COMPUTING↗

Inverse Modeling of Hydrologic Parameters in CLM4 via Generalized Polynomial Chaos in the Bayesian Framework

In this work, generalized polynomial chaos (gPC) expansion for land surface model parameter estimation is evaluated. We perform inverse modeling and compute the posterior distribution of the critical hydrological parameters that are subject to great uncertainty in the Community Land Model (CLM) for a given value of the output LH. The unknown parameters include those that have been identified as the most influential factors on the simulations of surface and subsurface runoff, latent and sensible heat fluxes, and soil moisture in CLM4.0. We set up the inversion problem in the Bayesian framework in two steps: (i) building a surrogate model expressing the input–output mapping, and (ii) performing inverse modeling and computing the posterior distributions of the input parameters using observation data for a given value of the output LH. The development of the surrogate model is carried out with a Bayesian procedure based on the variable selection methods that use gPC expansions. Our approach accounts for bases selection uncertainty and quantifies the importance of the gPC terms, and, hence, all of the input parameters, via the associated posterior probabilities.

97 MATHEMATICS AND COMPUTING↗

ADAPTIVE GROUP TESTING WITH MISMATCHED MODELS

Accurate detection of infected individuals is one of the critical steps in stopping any pandemic. When the underlying infection rate of the disease is low, testing people in groups, instead of testing each individual in the population, can be more efficient. In this work, we consider noisy adaptive group testing design with specific test sensitivity and specificity that select the optimal group given previous test results based on pre-selected utility function. As in prior studies on group testing, we model this problem as a sequential Bayesian Optimal Experimental Design (BOED) to adaptively design the groups for each test. We analyze the required number of group tests when using the updated posterior on the infection status and the corresponding Mutual Information (MI) as our utility function for selecting new groups. More importantly, we study how the potential bias on the ground-truth noise of group tests may affect the group testing sample complexity.

97 MATHEMATICS AND COMPUTING↗

Probabilistic Context Neighborhood model for lattices

Here we present the Probabilistic Context Neighborhood model designed for two-dimensional lattices as a variation of a Markov random field assuming discrete values. In this model, the neighborhood structure has a fixed geometry but a variable order, depending on the neighbors’ values. Our model extends the Probabilistic Context Tree model, originally applicable to one-dimensional space. It retains advantageous properties, such as representing the dependence neighborhood structure as a graph in a tree format, facilitating an understanding of model complexity. Furthermore, we adapt the algorithm used to estimate the Probabilistic Context Tree to estimate the parameters of the proposed model. We illustrate the accuracy of our estimation methodology through simulation studies. Additionally, we apply the Probabilistic Context Neighborhood model to spatial real-world data, showcasing its practical utility.

97 MATHEMATICS AND COMPUTING↗

Sequential Selection for Minimizing the Variance with Application to Crystallization Experiments

For many crystal-based products (e.g., pharmaceuticals, energy storage), the size uniformity is not only a key quality attribute, but sometimes also an indicator of other attributes such as solid purity. This article proposes a sequential selection approach to find a proper experimental setting that leads to high uniformity, or equivalently, small variance for crystal sizes, from the advanced slug flow reaction crystallization process of a model crystal, called manganese oxalate hydrate. The proposed sequential selection approach contains a Bayesian adaptive method to incorporate new uniformity measurements in each step and two design acquisition functions to improve the selection of the most promising experimental setting in terms of minimizing the variance. We study the performance of the proposed approach through multiple synthetic numerical studies, as well as a case study based on data from slug flow crystallization experiments. Throughout these studies, the proposed approach shows competitive performance in identifying the best experimental setting.

Expected improvement↗

Inverse Reinforcement Learning based Bayesian Goal Inference Method for Early Nuclear Proliferation Detection

Traditional methods for detection of nuclear proliferation indicators are usually applied after nuclear proliferation has already occurred. There is a need to advance these methods to perform early detection of nuclear proliferation indicators. In this project, we formulated an early detection problem as a sequential, decision-making, goal inference problem based on research publications of authors, to determine whether it is possible to infer whether an author will publish on a research activity before it has occurred. To develop and test our approach, we selected a civil nuclear activity for our case study. We constructed a state-action-state transition graph from publications of authors associated with the activity and the co-authors of their publications, using titles, abstracts, and author publication sequences. We then used inverse reinforcement learning to model the goal-directed behavior of authors in trajectories that terminate at selected goal states. Using a Bayesian formulation, we computed the probability that authors would reach each selected state from partially observed trajectories of their state transitions in their research topic space. The state with the highest probability was selected as the most probable goal state. Based on our results, we found that 60% of the times we can infer the correct goal state early; sometimes the inference is either delayed, or multiple states could be inferred as goal states. Overall, our results show that it is possible to perform early detection of research activities of authors in a nuclear technology area. Further research is necessary to establish a more accurate understanding of how topic modeling, topic space grid discretization, and the extent of overlap among trajectories of different goal states, affect the goal inference results. The methods developed in this work may be used to enhance data-driven methods for early detection of nuclear proliferation indicators.

97 MATHEMATICS AND COMPUTING↗

Capabilities of multivariate Bayesian inference toward seismic hazard assessment

Multivariate Bayesian analysis can bring significant benefits to seismic hazard analysis: Its multivariate feature enables computing scalar and vector hazard without making any approximations; Correlations between intensity measures are implicitly modeled, permitting direct simulation of ground motion selection tools such as the conditional mean spectrum and the generalized conditioning intensity measure; and Its updating feature enables a seamless integration of new ground motion data into the hazard results. Here, we first develop a multivariate Bayesian ground motion model through the NGA-West2 database. The model functional form considers fault-type, magnitude, and distance dependencies, and also the linear and the rock intensity dependent site response. We use a hybrid Markov Chain Monte Carlo sampling to perform Bayesian inference consisting of Gibbs step and a multilevel Metropolis-Hastings step. We then perform several checks on the model and note that its performance is satisfactory. Finally, we illustrate the merits of this multivariate Bayesian analysis, which include: ground motion model updating with ground motion data recorded in the last four years not part of the NGA-West2 database; computation of scalar and vector seismic hazard using the un-updated and updated ground motion models for Los Angeles, CA; and simulation of the conditional mean spectrum under scalar and vector IM conditioning while accounting for different sources of aleatoric and epistemic uncertainties.

58 GEOSCIENCES↗

Union through UNITY: Cosmology with 2000 SNe Using a Unified Bayesian Framework

Type Ia supernovae (SNe Ia) were instrumental in establishing the acceleration of the Universe’s expansion. By virtue of their combination of distance reach, precision, and prevalence, they continue to provide key cosmological constraints, complementing other cosmological probes. Individual SN surveys cover only over about a factor of 2 in redshift, so compilations of multiple SN data sets are strongly beneficial. We assemble an up-to-date “Union” compilation of 2087 cosmologically useful SNe Ia from 24 data sets (“Union3”). We take care to put all SNe on the same distance scale and update the light-curve fitting with SALT3 to use the full rest-frame optical. Over the next few years, the number of cosmologically useful SNe Ia will increase by more than a factor of 10, and keeping systematic uncertainties subdominant will be more challenging than ever. We discuss the importance of treating outliers, selection effects, light-curve shape/color populations/standardization relations, unexplained dispersion, and heterogeneous observations simultaneously. We present an updated Bayesian framework, called UNITY1.5 (Unified Nonlinear Inference for Type-Ia cosmologY), that incorporates significant improvements in our ability to model selection effects, standardization, and systematic uncertainties compared to earlier analyses. As an analysis byproduct, we also recover the posterior of the SN-only peculiar-velocity field, although we do not interpret it in this work. We compute updated cosmological constraints with Union3 and UNITY1.5, finding weak 1.7σ–2.6σ tension with flat cold dark matter and possible evidence for thawing dark energy (w0 > − 1, wa < 0). We release our SN distances, light-curve fits, and UNITY1.5 framework to the community.

Rubin, David↗

A Framework for Parametric and Predictive Uncertainty Quantification in the E3SM Land Model: Assessing Site and Observable Generalizability

Quantifying parametric uncertainty using observations from individual sites provides a critical foundation for Earth system modeling, serving as a necessary first step before scaling up to regional or global applications. This study introduces a novel computational framework designed to enhance model predictability by reducing parametric uncertainty and assessing site and observable generalizability using various observational constraints. The framework integrates five components: Model Simulation, Statistical Emulation, Global Sensitivity Analysis (GSA), Model Calibration, and Model Prediction. Using the E3SM land model, we simulated site-level land-atmosphere carbon and energy fluxes from 2003 to 2007 across five evergreen needleleaf FLUXNET sites, perturbing 26 vegetation-related model parameters. Gaussian process emulators were employed to expedite GSA and model calibration. Four critical parameters that strongly influence selected land-atmosphere fluxes were identified by GSA. Bayesian approaches were used to infer parameter probability distributions leveraging synthetic data and FLUXNET observations. The results reveal that posterior parameter distributions vary significantly across different sites and observables within the same plant functional type. Probabilistic predictions indicate that parameters calibrated at one site can enhance predictive accuracy at other sites, although site heterogeneity may sometimes outweigh parametric uncertainty. Additionally, the probabilistic predictions demonstrate that calibration for one variable can also improve predictability for other variables, thereby maximizing predictive capabilities with limited observations. This framework provides a powerful approach for reducing parametric uncertainty in Earth system models and deepening our understanding of carbon dynamics and energy cycles. Its adaptability makes it a valuable tool for broader applications in Earth system modeling.

54 ENVIRONMENTAL SCIENCES↗

Quantifying uncertainty for deep learning based forecasting and flow-reconstruction using neural architecture search ensembles

Classical problems in computational physics such as data-driven forecasting and signal reconstruction from sparse sensors have recently seen an explosion in deep neural network (DNN) based algorithmic approaches. However, most DNN models do not provide uncertainty estimates, which are crucial for establishing the trustworthiness of these techniques in downstream decision making tasks and scenarios. In recent years, ensemble-based methods have achieved significant success for the uncertainty quantification in DNNs on a number of benchmark problems. However, their performance on real-world applications remains under-explored. In this work, we present an automated approach to DNN discovery and demonstrate how this may also be utilized for ensemble-based uncertainty quantification. Specifically, we propose the use of a scalable neural and hyperparameter architecture search for discovering an ensemble of DNN models for complex dynamical systems. We highlight how the proposed method not only discovers high-performing neural network ensembles for our tasks, but also quantifies uncertainty seamlessly. This is achieved by using genetic algorithms and Bayesian optimization for sampling the search space of neural network architectures and hyperparameters. Subsequently, a model selection approach is used to identify candidate models for an ensemble set construction. Afterwards, a variance decomposition approach is used to estimate the uncertainty of the predictions from the ensemble. We demonstrate the feasibility of this framework for two tasks — forecasting from historical data and flow reconstruction from sparse sensors for the sea-surface temperature. In conclusion, we demonstrate superior performance from the ensemble in contrast with individual high-performing models and other benchmarks.

Deep ensembles↗

Bayesian inference for plasmonic nanometrology

Here, we introduce a Bayesian method for the characterization of plasmonic nanoparticles, which is applicable to both near- and far-field problems. Designed to combine data generated from any photon-plasmon interaction experiment with physically motivated theoretical models, our approach leverages state-of-the-art Markov chain Monte Carlo sampling techniques and returns parameter estimates on nanometric scales. Simulated spectral data sets, describing resonant scattering of photons from ellipsoidal and toroidal nanoparticles, are explored as concrete examples of our approach, with the resulting Bayesian estimates showing excellent agreement with the ground truth, even under conditions of high statistical noise. By incorporating Bayes factors into the method as well, we reveal how model selection can determine which one of competing geometric shapes better explains the observed data. Our comprehensive nanometrology procedure can be tailored to a variety of light-particle interaction models, and its reliance on Bayesian inference furnishes automatic uncertainty quantification. In addition to applicability to a host of plasmonic configurations such as nanoparticle dimers, trimers, and array studies, it is proposed that the presented analysis can be extended to the quantum regime, where nonclassical photon statistics may provide additional insight for inference of scatterer properties.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

PROVABGS: The Probabilistic Stellar Mass Function of the BGS One-percent Survey

We present the probabilistic stellar mass function (pSMF) of galaxies in the DESI Bright Galaxy Survey (BGS), observed during the One-percent Survey. The One-percent Survey was one of DESI's survey validation programs conducted from 2021 April to May, before the start of the main survey. It used the same target selection and similar observing strategy as the main survey and successfully observed the spectra and redshifts of 143,017 galaxies in the r < 19.5 magnitude-limited BGS Bright sample and 95,499 galaxies in the fainter surface-brightness- and color-selected BGS Faint sample over z < 0.6. We derive pSMFs from posteriors of stellar mass, M*, inferred from DESI photometry and spectroscopy using the Hahn et al. PRObabilistic Value-Added BGS (PROVABGS) Bayesian spectral energy distribution modeling framework. We use a hierarchical population inference framework that statistically and rigorously propagates the M* uncertainties. Furthermore, we include correction weights that account for the selection effects and incompleteness of the BGS observations. We present the redshift evolution of the pSMF in BGS, as well as the pSMFs of star-forming and quiescent galaxies classified using average specific star formation rates from PROVABGS. Overall, the pSMFs show good agreement with previous stellar mass function measurements in the literature. Our pSMFs showcase the potential and statistical power of BGS, which in its main survey will observe >100 × more galaxies. Moreover, we present the statistical framework for subsequent population statistics measurements using BGS, which will characterize the global galaxy population and scaling relations at low redshifts with unprecedented precision.

79 ASTRONOMY AND ASTROPHYSICS↗