Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Bayesian reasoning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

py-boomer v0.1.0

Py-BOOMER (Python Bayesian OWL Ontology MErgER in Python) is a probabilistic reasoning system for knowledge representation and ontological reasoning with uncertainty. Itnables reasoning over probabilistic facts and taxonomic relationships, finding the most likely consistent interpretation of potentially conflicting assertions. It uses a combination of graph-based reasoning and Bayesian probabilistic inference. Key features: Represent probabilistic ontological statements Reason over class subsumption hierarchies Evaluate class equivalence relationships Detect and resolve logical inconsistencies Calculate posterior probabilities for each assertion

Mungall, Chris [Lawrence Berkeley National Laborat

High-significance detection of correlation between the unresolved gamma-ray background and the large-scale cosmic structure

Our understanding of the γ-ray sky has improved dramatically in the past decade, however, the unresolved γ-ray background (UGRB) still has a potential wealth of information about the faintest γ-ray sources pervading the Universe. Statistical cross-correlations with tracers of cosmic structure can indirectly identify the populations that most characterize the γ-ray background. In this study, we analyze the angular correlation between the γ-ray background and the matter distribution in the Universe as traced by gravitational lensing, leveraging more than a decade of observations from the Fermi-Large Area Telescope (LAT) and 3 years of data from the Dark Energy Survey (DES). We detect a correlation at signal-to-noise ratio of 8.9. Most of the statistical significance comes from large scales, demonstrating, for the first time, that a substantial portion of the UGRB aligns with the mass clustering of the Universe as traced by weak lensing. Blazars provide a plausible explanation for this signal, especially if those contributing to the correlation reside in halos of large mass (∼ 10 14 M ⊙ ) and account for approximately 30–40% of the UGRB above 10 GeV. Additionally, we observe a preference for a curved γ-ray energy spectrum, with a log-parabolic shape being favored over a power-law. We also discuss the possibility of modifications to the blazar model and the inclusion of additional γ-ray sources, such as star-forming galaxies, misalinged active galactic nuclei, or particle dark matter.

79 ASTRONOMY AND ASTROPHYSICS

Galaxy cluster profiles: a Gaussian mixture model approach to halo miscentering

Measurements of the galaxy density and weak-lensing profiles of galaxy clusters typically rely on an assumed cluster center, which is taken to be the brightest cluster galaxy or other proxies for the true halo center defined as the minimum in the potential well. Departure of the assumed cluster center from the true halo center bias the resultant profile measurements, an effect known as miscentering bias. Currently, miscentering is typically modeled in stacked profiles of clusters with a two parameter model. We use an alternate approach in which the profiles of individual clusters are used with the corresponding likelihood computed using a Gaussian mixture model. We test the approach using halos and the corresponding subhalo profiles from the IllustrisTNG hydrodynamic simulations. We obtain significantly improved estimates of the miscentering parameters for both 3D and projected 2D profiles relevant for imaging surveys. We discuss applications to upcoming cosmological surveys. Our Python package for the Gaussian mixture model is publicly available at https://github.com/KyleMiller1/Halo-Miscentering-Mixture-Model.

Bayesian reasoning

Enhancing DESI DR1 full-shape analyses using HOD-informed priors

We present an analysis of DESI Data Release 1 (DR1) that incorporates Halo Occupation Distribution (HOD)-informed priors into Full-Shape (FS) modeling of the power spectrum based on cosmological perturbation theory (PT). By leveraging physical insights from the galaxy-halo connection, these HOD-informed priors on nuisance parameters substantially mitigate projection effects in extended cosmological models that allow for dynamical dark energy. The resulting credible intervals now encompass the posterior maximum from the baseline analysis using gaussian priors, eliminating a significant posterior shift observed in baseline studies. In the ΛCDM framework, a combined DESI DR1 FS information and constraints from the DESI DR1 baryon acoustic oscillations (BAO) — including Big Bang Nucleosynthesis (BBN) constraints and a weak prior on the scalar spectral index — yields Ω m = 0.2994 ± 0.0090 and σ 8 = 0.836$^{+0.024}_{-0.027}$, representing improvements of approximately 4% and 23% over the baseline analysis, respectively. For the w 0 w a CDM model, our results from various data combinations are highly consistent, with all configurations converging to a region with w 0 > -1 and w a < 0. This convergence not only suggests intriguing hints of dynamical dark energy but also underscores the robustness of our HOD-informed prior approach in delivering reliable cosmological constraints.

59 BASIC BIOLOGICAL SCIENCES

Frequentist cosmological constraints from full-shape clustering measurements in DESI DR1

We present a frequentist analysis of clustering measurements from Data Release 1 of the Dark Energy Spectroscopic Instrument (DESI) using the standard profile likelihood method. While Bayesian inferences for effective field theory models of galaxy clustering can be highly sensitive to prior choices for extended cosmological models, frequentist inferences are not susceptible to such effects. We compare frequentist and Bayesian constraints for the parameter set {σ 8 , H 0 , Ω m , w 0 , w a } using the full-shape power spectrum multipoles, post-reconstruction baryon acoustic oscillation (BAO) measurements, and external datasets from the CMB and type Ia supernovae measurements. The frequentist confidence intervals are significantly shifted relative to the Bayesian credible intervals for the w 0 w a CDM model, unless supernovae data are included. When DESI full-shape and BAO data are fit jointly, we obtain the following 1σ frequentist confidence intervals for ΛCDM (w 0 w a CDM): σ 8 = 0.863 +0.048 -0.040 , H 0 = 68.96 +0.81 -0.80 km s -1 Mpc -1 , Ω m = 0.3034 ± 0.0110 (σ 8 = 0.782 +0.060 -0.036 , H 0 = 63.7 +4.2 -2.0 km s -1 Mpc -1 , Ω m = 0.378 +0.024 -0.047 , w 0 = -0.16 +0.10 -0.50 , w a = -3.0 +1.7 ), corresponding to 0.8σ, 0.3σ, 0.7σ (2.1σ, 4.1σ, 6.5σ, 6.3σ, 6.6σ) shifts between the maximum likelihood estimate and the Bayesian posterior mean for ΛCDM (w 0 w a CDM) respectively.

Bayesian reasoning

Testing T2K’s Bayesian constraints with priors in alternate parameterisations

Bayesian analysis results require a choice of prior distribution. In long-baseline neutrino oscillation physics, the usual parameterisation of the mixing matrix induces a prior that privileges certain neutrino mass and flavour state symmetries. Here we study the effect of privileging alternate symmetries on the results of the T2K experiment. We find that constraints on the level of CP violation (as given by the Jarlskog invariant) are robust under the choices of prior considered in the analysis. On the other hand, the degree of octant preference for the atmospheric angle depends on which symmetry has been privileged.

Bayesian Inference

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES

Probabilistic inference in very large universes

Our current favored cosmological theories allow for the striking and controversial possibility that the observable universe is just a small part of a much larger universe in which parameters that describe the effective, low-energy laws of physics vary from one region to another. The controversy is largely driven by the fact that such a “very large universe” is mostly observationally inaccessible to us, so the issue arises of how we can reasonably assess a theory that describes such a universe. In this paper, we propose a Bayesian method for theory assessment based on theory-generated probability distributions for our observations. We focus on the principles that define this method, leaving aside concerns about how, in practice, one would carry out the required calculations. (One important issue that we set aside is the measure problem.) We argue that cosmological theories can be tested by the standard method of Bayesian updating, but we need to use theoretical predictions for “first-person” probabilities—that is, probabilities that we should use for our observations, taking into account all relevant selection effects. These selection effects can vary from one observer to another and can vary with time, so, in principle, first-person probabilities are defined for each observer instant—an observer at a specific instant of time. Calculations of first-person probabilities should take into account everything that the observer believes about herself and her surroundings, which we refer to as her subjective state. If the universe is very large, a theory might predict that there are many observer instants in the same subjective state; we argue that first-person probabilities should be calculated using a principle of self-locating indifference (PSLI), the assumption that any real observer should make predictions for her future as if she were chosen randomly and uniformly from the theoretically predicted observer instants that share her subjective state. We believe the PSLI is intuitively very reasonable, but we also argue that, if the theory is correct, the use of this principle maximizes the expected fraction of observers who will make correct predictions. A further complication is that cosmological theories are not expected to fully predict the detailed properties of the universe, but rather will predict a set of possible universes, each with a probability. Different possible universes will generically have different numbers of observers. We argue that, in the calculation of first-person probabilities, the probability for each possible universe should be weighted by the number of observer instants in the specified subjective state that it contains. These issues have been controversial in the literature, so we also provide a rebuttal to the claim that principles like the PSLI involve a “selection fallacy”; a rebuttal to what we dub the principle of required certainty; an argument rejecting theories that predict a preponderance of Boltzmann brains; a rebuttal to a parable about humans and Jovians used by Hartle and Srednicki to argue that assumptions of typicality can lead to absurd consequences; and, finally, a discussion about how the use of “old evidence” can be fit into a Bayesian mold.

Azhar, Feraz [University of Notre Dame, IN (United

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING

Modeling of Stress and Temperature Effects on Creep of Reduced Activation Ferritic-Martensitic Steel Alloy F82H (Tertiary Creep Modeling of RAFM Steel)

A Bayesian optimization procedure is presented for calibrating a multi-mechanism micromechanical model for creep to experimental data of F82H steel. Reduced activation ferritic martensitic (RAFM) steels based on are the most promising candidates for some fusion reactor structures. Although there are indications that RAFM steel could be viable for fusion applications at temperatures up to 600 °C, the maximum operating temperature will be determined by the creep properties of the structural material and the breeder material compatibility with the structural material. Due to the relative paucity of available creep data on F82H steel compared to other alloys such as Grade 91 steel, micromechanical models are sought for simulating creep based on relevant deformation mechanisms. As a point of departure, this work recalibrates a model form that was previously proposed for Grade 91 steel to match creep curves for F82H steel. Due to the large number of parameters (9) and cost of the nonlinear simulations, an automated approach for tuning the parameters is pursued using a recently developed Bayesian optimization for functional output (BOFO) framework [1]. Incorporating extensions such as batch sequencing and weighted experimental load cases into BOFO, a reasonably small error between experimental and simulated creep curves at two load levels is achieved in a reasonable number of iterations. Validation with an additional creep curve provides confidence in the fitted parameters obtained from the automated calibration procedure to describe the creep behavior of F82H steel at 600 °C. The model is further extended using a temperature dependent scaling law approach to simulate creep response between 550 °C and 650 °C. The efficacy of this extension is compared with the previously used scaling law approach for Grade 91 steel.

36 MATERIALS SCIENCE

Calibration of RAFM Micromechanical Model for Creep Using Bayesian Optimization for Functional Output

A Bayesian optimization procedure is presented for calibrating a multimechanism micromechanical model for creep to experimental data of F82H steel. Reduced activation ferritic martensitic (RAFM) steels based on Fe(8–9)%Cr are the most promising candidates for some fusion reactor structures. Although there are indications that RAFM steel could be viable for fusion applications at temperatures up to 600°C, the maximum operating temperature will be determined by the creep properties of the structural material and the breeder material compatibility with the structural material. Due to the relative paucity of available creep data on F82H steel compared to other alloys such as Grade 91 steel, micromechanical models are sought for simulating creep based on relevant deformation mechanisms. As a point of departure, this work recalibrates a model form that was previously proposed for Grade 91 steel to match creep curves for F82H steel. Due to the large number of parameters (9) and cost of the nonlinear simulations, an automated approach for tuning the parameters is pursued using a recently developed Bayesian optimization for functional output (BOFO) framework (Huang et al., 2021, “Bayesian optimization of functional output in inverse problems,” Optim. Eng., 22, pp. 2553–2574). Incorporating extensions such as batch sequencing and weighted experimental load cases into BOFO, a reasonably small error between experimental and simulated creep curves at two load levels is achieved in a reasonable number of iterations. In conclusion, validation with an additional creep curve provides confidence in the fitted parameters obtained from the automated calibration procedure to describe the creep behavior of F82H steel.

42 ENGINEERING

Automated and highly parallelized Bayesian optimization scheme for direct drive fusion experiments on OMEGA

Finding the optimal implosion design on existing experimental facilities for inertial confinement fusion requires an exhaustive search of the vast design parameter space. This is infeasible both with experiments and with simulations. Consequently, a large fraction of the experimentally realizable design space remains unexplored, and new design schemes are challenging to optimize in a reasonable time frame. On the OMEGA laser facility, predictive machine learning models have been developed to accurately forecast the result of an experiment using only inexpensive simulations and the large dataset of prior experimental data. However, the full design space remains vast enough to be unassailable with simple optimization techniques. Here we develop an automated and optimally parallel Bayesian optimization algorithm that can entirely optimize the target and pulse shape of a direct-drive ICF implosion under a given design paradigm. We use this algorithm to find a markedly improved design for the performance implosions on OMEGA that is predicted to hydroequivalently scale to ignition at 2.15 MJ.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Propagating synthetic populations with dynamic Bayesian networks: a framework for long-horizon demographic forecasting

This study presents a dynamic demographic microsimulator using dynamic Bayesian networks to forecast long–term changes in household and individual life events. Leveraging longitudinal Panel Study of Income Dynamics (PSID) data, two networks for individuals and households were modeled to simulate transitions in employment, income, education, marriage, childbirth, leaving the parental home, home ownership, mortality, and household formation or dissolution. Across 1,000 simulation runs spanning 24 years, household–level outcomes remain highly accurate and individual–level predictions reasonable. Although accuracy naturally declines with projection horizon, performance remains promising at both levels. This study addresses a key limitation of existing population synthesis models, which typically generate only a single static snapshot of the population. In conclusion, by introducing a framework that propagates cross-sectional outputs into the future, the microsimulator enables the tracking of demographic evolution over time, enhances realism in population-based simulations, and supplies credible inputs to agent-based travel demand models.

Demographic modeling

Mechanistic within-host mathematical model of inhalational anthrax

We present a mathematical model of the dynamics of Bacillus anthracis bacteria within the lymph nodes and blood of a host, following inhalation of an initial dose of spores. We also incorporate the dynamics of protective antigen, which is the binding component of the anthrax toxin produced by the bacteria. The model offers a mechanistic description of the early infection dynamics of inhalational anthrax, while its stochastic nature allows us to study the probabilities of different outcomes (for example, how likely it is that the infection will be cleared for a given inhaled dose of spores) in order to explain dose-response data for inhalational anthrax. The model is calibrated via a Bayesian approach, using in vivo data from New Zealand white rabbit and guinea pig infection studies, enabling within-host parameters to be estimated. We also leverage incubation-period data from the Sverdlovsk 1979 anthrax outbreak to show that the model can accurately describe human time-to-symptoms data under reasonable parameter regimes. Finally, we derive a simple approximate formula for the probability of symptom onset before time t, assuming that the number of inhaled spores has a Poisson distribution.

59 BASIC BIOLOGICAL SCIENCES

Rapid Bayesian High Entropy Alloy Designs Fabricated via Wire Arc Additive Manufacturing

Purpose: This project seeks to demonstrate a new high-throughput (rapid) alloy design technique applied to creating new high entropy alloys (HEAs) for extreme environments. High entropy alloys shift the design paradigm from being focused on a single principal element (e.g. nickel-based alloys) to target alloys that include high atomic fractions (X >10%) of multiple elements. These HEA materials can exhibit sluggish diffusion and enhanced corrosion resistance, ideal for potential applications in advanced ultra supercritical (A-USC) steam cycles for power generation. Scope: The addition of multiple elements in high atomic fractions creates an enormous design space that cannot easily be investigated by traditional material design strategies such as designed of experiments (DOE). This project utilizes a Bayesian machine learning algorithm that has been modified to work with calculation of phase diagrams (CALPHAD) software. This Bayesian algorithm reduces manual inputs and increase the likelihood of achieving an optimal solution. Compositional inputs to this algorithm will be assessed using existing material property models for high temperature strength and corrosion resistance. The target for alloy performance will be a 15% (~100 ⁰C) increase in allowable service temperature beyond heat-resistant stainless steels while maintaining or improving alloy cost and corrosion resistance. Haynes 230 was selected as a baseline, which is 57 wt% Ni with 22 wt% Cr 14 wt% W, and 2 wt% Mo as solid solution strengtheners. In addition to rapid design via Bayesian machine learning, the alloys were rapidly fabricated using a multi-wire arc additive manufacturing (mWAAM) technique which allows for precise control of alloy composition and assessing of alloy design “windows” to study composition effects. Build speeds for wire-arc additive processes are among the highest for additive technologies enabling rapid and reliable sample fabrication when compared to conventional methods such as arc button melting. The mWAAM samples will be rapidly characterized via instrumented indentation for room temperature modulus and strength and for elevated temperature strength via hot hardness tests. After being screened with hardness testing, potential alloys will be further evaluated with conventional microscopy techniques including scanning electron microscopy (SEM) and transmission electron microscopy (TEM) to assess agreement with modeling results. The most promising compositions will also be evaluated by printing full sized tensile specimens for mechanical behavior tests at elevated temperatures. Results: Bayesian machine learning of a single performance function was initially used to optimize five performance metrics: 1) single phase stability, 2) yield strength, 3) creep resistance (low diffusion coefficient), 4) freezing range (weldability), and 5) material cost. The single performance function was suboptimal as assumptions had to be made about the results while formulating the optimization. A goal-oriented Bayesian optimization strategy (Hanaoka, 2021) was implemented with CALPHAD for use with the five metrics above. This multi-objective Bayesian optimization (MOBO) enabled the design of NiCrCoFe alloys with V and W additions. A base composition of NiCoCr was selected as Ni provides a stable FCC matrix, Cr aids corrosion/oxidation resistance, and Co is a solid-solutions strengthener that also improves creep by increasing the activation energy. Fe helps reduce diffusion coefficients and cost. Finally, V and W were selected for their reasonable solubility and high atomic misfit to aid in solid solution strengthening. Cracking of the mWAAM specimens was an early issue, and the Easton solidification cracking model (Easton et al., 2014a) was selected for addition to the MOBO function. High performing alloys fabricated by mWAAM included Ni 28 Cr 25 Co 26 Fe 15 V 8 and Ni 62 Cr 18 Co 1 Fe 3 W 15 . It was observed that even after adapting the mWAAM process for W, the W did not fully dissolve. To fully evaluate the Ni 62 Cr 18 Co 1 Fe 3 W 15 composition, a cored wire (80-20 NiCr sheath/powder core) was manufactured and printed via WAAM, and HIP’ing was utilized to homogenize and densify the printed alloy. The V and W alloys produced met metrics 1 (solid solution), 4 (solidification cracking), and 5 (cost). However, an unmodeled mechanism of thermal stress cracking was identified in the WAAM produced materials, perhaps exacerbated by the lack of grain boundary strengthening elements (B, C). Conclusions & Recommendations: A high-throughput (rapid) alloy design technique was applied to designing and manufacturing new high entropy alloys (HEAs) for extreme environments utilizing MOBO and mWAAM. The developed process was rapid and effective in addressing the mechanisms included in the model. The lack of grain boundary strengthening element additions (e.g., B, C) was a simplification that likely produced thermal stress cracking that turned into a large part of the investigation. Additions on the order of 0.005 wt% B and 0.05 wt% C likely would have minimized thermal stress grain boundary cracking. Overall, the high throughput design strategy is promising for rapid design of metrics-driven alloys for advanced ultra supercritical (A-USC) steam cycles for power generation. The MOBO and mWAAM process could be commercialized to accelerate metrics-driven alloy design. In addition, the cored-wire process utilized for scale-up is a promising high-volume process for WAAM alloy development and scale-up.

36 MATERIALS SCIENCE

Boron Coordination in Multicomponent Glasses: Analytical Models and Machine Learning With Uncertainty

Borosilicate glasses are extensively used in a variety of applications from kitchenware to nuclear waste immobilization due to the strong network formed by the Si-O-B bond that makes it resistant to chemical corrosion and gives it a low thermal expansion. Boron, however, exists in both trigonal BO3 and tetrahedral BO4 bonds in glass systems, which impacts the chemical durability and thermal resistance of the glass, amongst other properties. Boron coordination (N4), or the ratio of the amount of BO4 to BO3 within a glass, may aid in predicting these properties but is difficult to derive without experimental data due to the complexity of impacts from varied glass compositions and processing factors. For this reason, compositional models have been developed to predict boron coordination, but the models typically include a limited number of glass components. To help fill this gap in the models, in this work, a diverse multicomponent glass dataset of 809 glasses is compiled from a literature search, and then a number of analytical and machine learning (ML) models are trained on the dataset. Previously developed modified Bernstein and modified Du Stebbins analytical models were fitted to update parameters with the new dataset. Then, partially Bayesian neural networks, Gaussian process regressor, and heteroskedastic deterministic neural networks were evaluated. The ML models examined all have different strategies to overcome the potential for overfitting as a result of a limited training dataset, and return results that account for model uncertainty, which can be valuable for understanding model reliability. For the first time, cooling rate is introduced as an input parameter for ML models, showing consistent improvements in performance and solidifying the importance of including parameters outside of composition alone for N4 prediction. The machine learning models examined here show promise in accurate predictions of boron coordination in borosilicate glasses, all achieving R2 values of 0.91.

boron coordination

Sharp detection of low-dimensional structure in probability measures via dimensional logarithmic Sobolev inequalities

Identifying low-dimensional structure in high-dimensional probability measures is an essential pre-processing step for efficient sampling. To identify this structure, we approximate the target measure as a perturbation of an arbitrary reference measure along a few directions in $\mathbb{R}^{d}$. These directions are determined by minimizing an upper bound on the Kullback–Leibler (KL) divergence between the target and its approximation. Our contribution improves upon previous works by leveraging dimensional logarithmic Sobolev inequalities to refine the bound on the KL divergence. These inequalities lead to a uniformly tighter bound on the KL divergence, thereby enhancing the identification of the most significant perturbation directions. In particular, when the target and reference are both Gaussian, minimizing the resulting bound is equivalent to minimizing the KL divergence. We further demonstrate the applicability of this analysis to the squared Hellinger distance, where analogous reasoning shows that the dimensional Poincaré inequality offers improved bounds.

Bayesian inference

Evaluating the limitations of Bayesian metabolic control analysis

Bayesian Metabolic Control Analysis (BMCA) is a promising framework for inferring metabolic control coefficients in data-limited scenarios, combining Bayesian inference with linear-logarithmic (lin-log) rate laws. These metabolic control coefficients quantify how changes in enzyme activities affect steady-state fluxes and metabolite concentrations across a metabolic network. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCC), and concentration control coefficients (CCC) under varying data availability conditions using three synthetic metabolic network models. We demonstrate that BMCA predictions are highly dependent on the inclusion of flux and enzyme concentration data, with the omission of these datasets leading to severe inaccuracies. In our synthetic, enzyme-perturbation datasets, external metabolite concentrations had minimal impact and, in some cases, their exclusion improved predictions; when external-nutrient perturbations were introduced and those concentrations were observed, gains were at most modest. Additionally, we find that posterior estimation with both ADVI and HMC can underestimate large-magnitude elasticities in our synthetic settings, with ADVI showing somewhat higher variance under strong up-regulation; thus, recovering |elasticity| ≳ 1.5 remains challenging regardless of the inference engine. ADVI also fails to accurately infer allosteric interactions, even when regulatory effects are strong. While BMCA maintains reasonable accuracy in partially recovering the rankings of the highest FCC values, its estimates of absolute values remain constrained by prior assumptions and data limitations. Our findings reveal the BMCA algorithm’s strengths and weaknesses, providing guidance on its application in metabolic engineering, and highlighting the need for methodological refinements to enhance its predictive capabilities.

59 BASIC BIOLOGICAL SCIENCES