Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The NANOGrav Nine-Year Data Set: Limits on the Isotropic Stochastic Gravitational Wave Background

We compute upper limits on the nanohertz-frequency isotropic stochastic gravitational wave background (GWB) using the 9 year data set from the North American Nanohertz Observatory for Gravitational Waves (NANOGrav) collaboration. Well-tested Bayesian techniques are used to set upper limits on the dimensionless strain amplitude (at a frequency of 1 yr(exp -1) for a GWB from supermassive black hole binaries of A(sub gw) less than 1.5 x 10(exp -15). We also parameterize the GWB spectrum with a broken power-law model by placing priors on the strain amplitude derived from simulations of Sesana and McWilliams et al. Using Bayesian model selection we find that the data favor a broken power law to a pure power law with odds ratios of 2.2 and 22 to one for the Sesana and McWilliams prior models, respectively. Using the broken power-law analysis we construct posterior distributions on environmental factors that drive the binary to the GW-driven regime including the stellar mass density for stellar-scattering, mass accretion rate for circumbinary disk interaction, and orbital eccentricity for eccentric binaries, marking the first time that the shape of the GWB spectrum has been used to make astrophysical inferences. Returning to a power-law model, we place stringent limits on the energy density of relic GWs, OMEGA(sub gw) (f) h squared less than 4.2 x 10(exp -10). Our limit on the cosmic string GWB, OMEGA(sub gw) (f) h squared less than 2.2 x 10(exp -10), translates to a conservative limit on the cosmic string tension with G mu less than 3.3 x 10(exp -8), a factor of four better than the joint Planck and high-l‚ cosmic microwave background data from other experiments.

Arzoumanian, Z.↗

AT2017gfo: Bayesian inference and model selection of multicomponent kilonovae and constraints on the neutron star equation of state

The joint detection of the gravitational wave GW170817, of the short γ-ray burst GRB170817A and of the kilonova AT2017gfo, generated by the the binary neutron star (NS) merger observed on 2017 August 17, is a milestone in multimessenger astronomy and provides new constraints on the NS equation of state. We perform Bayesian inference and model selection on AT2017gfo using semi-analytical, multicomponents models that also account for non-spherical ejecta. Observational data favour anisotropic geometries to spherically symmetric profiles, with a log-Bayes’ factor of ~104, and favour multicomponent models against single-component ones. The best-fitting model is an anisotropic three-component composed of dynamical ejecta plus neutrino and viscous winds. Using the dynamical ejecta parameters inferred from the best-fitting model and numerical–relativity relations connecting the ejecta properties to the binary properties, we constrain the binary mass ratio to q < 1.54 and the reduced tidal parameter to $120\lt \tilde{\Lambda }\lt 1110$. Finally, we combine the predictions from AT2017gfo with those from GW170817, constraining the radius of a NS of 1.4 M ⊙ to 12.2 ± 0.5 km (1σ level). This prediction could be further strengthened by improving kilonova models with numerical-relativity information.

79 ASTRONOMY AND ASTROPHYSICS↗

A Bayesian framework for adsorption energy prediction on bimetallic alloy catalysts

Abstract For high-throughput screening of materials for heterogeneous catalysis, scaling relations provides an efficient scheme to estimate the chemisorption energies of hydrogenated species. However, conditioning on a single descriptor ignores the model uncertainty and leads to suboptimal prediction of the chemisorption energy. In this article, we extend the single descriptor linear scaling relation to a multi-descriptor linear regression models to leverage the correlation between adsorption energy of any two pair of adsorbates. With a large dataset, we use Bayesian Information Criteria (BIC) as the model evidence to select the best linear regression model. Furthermore, Gaussian Process Regression (GPR) based on the meaningful convolution of physical properties of the metal-adsorbate complex can be used to predict the baseline residual of the selected model. This integrated Bayesian model selection and Gaussian process regression, dubbed as residual learning, can achieve performance comparable to standard DFT error (0.1 eV) for most adsorbate system. For sparse and small datasets, we propose an ad hoc Bayesian Model Averaging (BMA) approach to make a robust prediction. With this Bayesian framework, we significantly reduce the model uncertainty and improve the prediction accuracy. The possibilities of the framework for high-throughput catalytic materials exploration in a realistic setting is illustrated using large and small sets of both dense and sparse simulated dataset generated from a public database of bimetallic alloys available in Catalysis-Hub.org.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Application of a Bayesian Framework for Plasticity Model Selection

Interpretable Machine Learning (IML) has performed well when tasked with deriving constitutive material models. However, IML has been shown to prefer models that overfit noise in data, which tends to lead to bloat and a decrease in interpretability. Due to these issues, the ability of IML to reliably derive models that fit the data and are both interpretable and generalizable is limited. A method developed recently has shown promise to improve upon traditional IML by using a Bayesian fitness definition for the evolution of free-form models with non-deterministic parameters. This framework was developed for genetic-programming-based symbolic regression(GPSR) and involves model parameter estimation using Sequential Monte Carlo sampling (SMC).The method has demonstrated a reduction in bloat when dealing with noisy data in comparison to conventional GPSR. The results of this framework applied to stress-strain data for copper show models that more effectively predict the experimental data better than was previously shown with GPSR.

plasticity↗

Hybrid Parameter Search and Dynamic Model Selection for Mixed-Variable Bayesian Optimization

Herein this article presents a new type of hybrid model for Bayesian optimization (BO) adept at managing mixed variables, encompassing both quantitative (continuous and integer) and qualitative (categorical) types. Our proposed new hybrid models (named hybridM) merge the Monte Carlo Tree Search structure (MCTS) for categorical variables with Gaussian Processes (GP) for continuous ones. hybridM leverages the upper confidence bound tree search (UCTS) for MCTS strategy, showcasing the tree architecture’s integration into Bayesian optimization. Our innovations, including dynamic online kernel selection in the surrogate modeling phase and a unique UCTS search strategy, position our hybrid models as an advancement in mixed-variable surrogate models. Numerical experiments underscore the superiority of hybrid models, highlighting their potential in Bayesian optimization.

97 MATHEMATICS AND COMPUTING↗

Investigation of Ethane Dehydrogenation and Hydrogenolysis on Pt(111), Pt(211), and Pt(100): Bayesian Quantification and Correction of DFT-Based Enthalpic and Entropic Uncertainties

Computational investigations of heterogeneously catalyzed reactions using density functional theory (DFT) are often inaccurate, largely due to uncertainties in the choice of DFT functional (enthalpic uncertainty) and approximations for modeling adsorbate movement along the catalyst surface (entropic uncertainty). This work illustrates that both uncertainties are significant in the investigation of ethane dehydrogenation (EDH) and hydrogenolysis on Pt catalysts by considering the complete deconstruction of ethane on Pt(111), Pt(211), and Pt(100) using microkinetic modeling (MKM). Hence, this work uses both noncalibrated and Bayesian-calibrated MKMs to quantify and correct inaccuracies in macroscopic properties due to both uncertainties. A Bayesian approach to the correction of entropic errors was introduced using a “Modified Fermi Function (MFF)” to calibrate between the two bounds of entropy represented by the harmonic oscillator (HO) and free translator (FT) approximations. Regardless of enthalpic and entropic uncertainties, all three surfaces are capable of ethane activation; however, Pt(211) was found to be the most active and is largely responsible for methane production. Next, Pt(111) is largely responsible for acetylene production, and Pt(100) has the highest ethylene selectivity but is most susceptible to coking. By comparison of different calibrated models, the FT entropy approximation was found to better describe EDH under typical experimental conditions. Statistical evidence was found to support Pt(111) as the active site for EDH, assuming that one single site is responsible for the chemistry. On the three surfaces, competing second dehydrogenations to CH 2 CH 2 and CH 3 CH were observed as well as isomerization of CH 3 CH back to CH 2 CH 2 and deeper dehydrogenation of CH 3 CH. In conclusion, C–C cleavage was found to largely proceed via the CH 3 C intermediate on Pt(100) and Pt(111), while on Pt(211), it was via both CHC and CH 3 C.

Bayesian model selection↗

Modeling the Effect of Surface Platinum–Tin Alloys on Propane Dehydrogenation on Platinum–Tin Catalysts

Uncertainty analysis, reported experimental literature data, and density functional theory were synthesized to model the effect of surface tin coverage on platinum-based catalysts for nonoxidative propane dehydrogenation to propylene. Here, this study tests four different platinum–tin skin surface models as potential catalytic sites, Pt 3 Sn/Pt(100), PtSn/Pt(100), Pt 3 Sn/Pt(111), and Pt 2 Sn/Pt(211), and compares them to the corresponding pure Pt surface sites using an uncertainty analysis methodology that uses BEEF-vdW with its ensembles (BMwE) to generate the uncertainty for the energies of the intermediates and transition states. One experimental data set with two experimental observations, selectivity to propylene and turnover frequency of propylene, was used as a calibration data set to evaluate the impact of the experimental data on informing the models. This study finds that the prior model for Pt 3 Sn/Pt(100) is the most active and Pt 2 Sn/Pt(211) is the most selective toward propylene. Active sites on the (100) facet have the highest probability of being responsible for C 1 and C 2 product formations (C–C bond cleavage). Increasing the Sn coverage on the (100) surface facet to a PtSn/Pt(100) active site leads to a significantly reduced rate and might explain the experimentally observed higher selectivity of Sn-doped catalysts relative to pure Pt catalysts. Next, this study finds that for all surfaces, except PtSn/Pt(100), the rate-controlling steps are the initial dehydrogenation steps alongside some partially rate-controlling second dehydrogenation steps. For PtSn/Pt(100), only the initial terminal dehydrogenation step to CH 3 CH 2 CH 2 * and second dehydrogenation steps are rate-controlling. Next, the calibrated models for all surfaces were found to be selective toward propylene production and model the reported turnover frequency successfully. Nevertheless, Pt 2 Sn/Pt(211) emerges as the active site with some (minor) evidence as the main active site based on Jeffreys’ scale interpretation of Bayes factors. This observation agrees with prior studies that also found step sites to be most likely the most relevant active sites for pure Pt catalysts. Overall, the results indicate that tin, in addition to affecting the binding strength of the adsorbed species, prevents deeper dehydrogenation (reducing coking) and cracking reactions through increasing activation barriers for unwanted side reactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nonlinear sparse Bayesian learning for physics-based models

This paper addresses the issue of overfitting while calibrating unknown parameters of over-parameterized physics-based models with noisy and incomplete observations. Here, a semi-analytical Bayesian framework of nonlinear sparse Bayesian learning (NSBL) is proposed to identify sparsity among model parameters during Bayesian inversion. NSBL offers significant advantages over machine learning algorithm of sparse Bayesian learning (SBL) for physics-based models, such as 1) the likelihood function or the posterior parameter distribution is not required to be Gaussian, and 2) prior parameter knowledge is incorporated into sparse learning (i.e. not all parameters are treated as questionable). NSBL employs the concept of automatic relevance determination (ARD) to facilitate sparsity among questionable parameters through parameterized prior distributions. The analytical tractability of NSBL is enabled by employing Gaussian ARD priors and by building a Gaussian mixture-model approximation of the posterior parameter distribution that excludes the contribution of ARD priors. Subsequently, type-II maximum likelihood is executed using Newton's method whereby the evidence and its gradient and Hessian information are computed in a semi-analytical fashion. We show numerically and analytically that SBL is a special case of NSBL for linear regression models. Subsequently, a linear regression example involving multimodality in both parameter posterior pdf and model evidence is considered to demonstrate the performance of NSBL in cases where SBL is inapplicable. Next, NSBL is applied to identify sparsity among the damping coefficients of a mass-spring-damper model of a shear building frame. These numerical studies demonstrate the robustness and efficiency of NSBL in alleviating overfitting during Bayesian inversion of nonlinear physics-based models.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantifying shifts in natural selection on codon usage between protein regions: a population genetics approach

Codon usage bias (CUB), the non-uniform usage of synonymous codons, occurs across all domains of life. Adaptive CUB is hypothesized to result from various selective pressures, including selection for efficient ribosome elongation, accurate translation, mRNA secondary structure, and/or protein folding. Given the critical link between protein folding and protein function, numerous studies have analyzed the relationship between codon usage and protein structure. The results from these studies have often been contradictory, likely reflecting the differing methods used for measuring codon usage and the failure to appropriately control for confounding factors, such as differences in amino acid usage between protein structures and changes in the frequency of different structures with gene expression. Here we take an explicit population genetics approach to quantify codon-specific shifts in natural selection related to protein structure in S. cerevisiae and E. coli. Unlike other metrics of codon usage, our approach explicitly separates the effects of natural selection, scaled by gene expression, and mutation bias while naturally accounting for a region’s amino acid usage. Bayesian model comparisons suggest selection on codon usage varies only slightly between helix, sheet, and coil secondary structures and, similarly, between structured and intrinsically-disordered regions. Similarly, in contrast to previous findings, we find selection on codon usage only varies slightly at the termini of helices in E. coli. Using simulated data, we show this previous work indicating “non-optimal” codons are enriched at the beginning of helices in S. cerevisiae was due to failure to control for various confounding factors (e.g. amino acid biases, gene expression, etc.), and rather than selection to modulate cotranslational folding. Our results reveal a weak relationship between codon usage and protein structure, indicating that differences in selection on codon usage between structures are slight. In addition to the magnitude of differences in selection between protein structures being slight, the observed shifts appear to be idiosyncratic and largely codon-specific rather than systematic reversals in the nature of selection. Overall, our work demonstrates the statistical power and benefits of studying selective shifts on codon usage or other genomic features from an explicitly evolutionary approach. Limitations of this approach and future potential research avenues are discussed.

59 BASIC BIOLOGICAL SCIENCES↗

Bayesian inference in band excitation scanning probe microscopy for optimal dynamic model selection in imaging

The universal tendency in scanning probe microscopy (SPM) over the last two decades is to transition from simple 2D imaging to complex detection and spectroscopic imaging modes. The emergence of complex SPM engines brings forth the challenge of reliable data interpretation, i.e., conversion from detected signals to descriptors specific to tip–surface interactions and subsequently to material’s properties. In this work, we implemented a Bayesian inference approach for the analysis of the image formation mechanisms in band excitation SPM. Compared to the point estimates in classical functional fit approaches, Bayesian inference allows for the incorporation of extant knowledge of materials and probe behavior in the form of corresponding prior distribution and return the information on the material functionality in the form of readily interpretable posterior distributions. We explore the nonlinear mechanical behaviors spatially in a classical ferroelectric material, PbTiO 3 . We observe the non-trivial evolution of the Duffing stiffness term and the nonlinearity of the sample surface, determine spatial clustering of the nonlinear response, and perform a Landau analysis on predicting the nonlinear coefficient, which indicates that ferroelectric behavior can be a cause of the observed results. These observations suggest that the spectrum of anomalous behaviors at the ferroelectric domain walls may be broader than previously believed and can extend to non-conventional mechanical properties in addition to static and microwave conductance.

36 MATERIALS SCIENCE↗

Wolf-Rayet Galaxies in SDSS-IV MaNGA. II. Metallicity Dependence of the High-mass Slope of the Stellar Initial Mass Function

As hosts of living high-mass stars, Wolf-Rayet (WR) regions or WR galaxies are ideal objects for constraining the high-mass end of the stellar initial mass function (IMF). We construct a large sample of 910 WR galaxies/regions that cover a wide range of stellar metallicity (from Z ~ 0.001 to 0.03) by combining three catalogs of WR galaxies/regions previously selected from the SDSS and SDSS-IV/MaNGA surveys. We measure the equivalent widths of the WR blue bump at ~4650 Å for each spectrum. They are compared with predictions from stellar evolutionary models Starburst99 and BPASS, with different IMF assumptions (high-mass slope α of the IMF ranging from 1.0 to 3.3). Both singular evolution and binary evolution are considered. We also use a Bayesian inference code to perform full spectral fitting to WR spectra with stellar population spectra from BPASS as fitting templates. We then make a model selection among different α assumptions based on Bayesian evidence. These analyses have consistently led to a positive correlation of the IMF high-mass slope α with stellar metallicity Z, i.e., with a steeper IMF (more bottom-heavy) at higher metallicities. Specifically, an IMF with α = 1.00 is preferred at the lowest metallicity (Z ~ 0.001), and an Salpeter or even steeper IMF is preferred at the highest metallicity (Z ~ 0.03). These conclusions hold even when binary population models are adopted.

79 ASTRONOMY AND ASTROPHYSICS↗