Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Frequentist statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

An implementation of neural simulation-based inference for parameter estimation in ATLAS

Neural simulation-based inference (NSBI) is a powerful class of machine-learning-based methods for statistical inference that naturally handles high-dimensional parameter estimation without the need to bin data into low-dimensional summary histograms. Such methods are promising for a range of measurements, including at the Large Hadron Collider, where no single observable may be optimal to scan over the entire theoretical phase space under consideration, or where binning data into histograms could result in a loss of sensitivity. This work develops a NSBI framework for statistical inference, using neural networks to estimate probability density ratios, which enables the application to a full-scale analysis. It incorporates a large number of systematic uncertainties, quantifies the uncertainty due to the finite number of events in training samples, develops a method to construct confidence intervals, and demonstrates a series of intermediate diagnostic checks that can be performed to validate the robustness of the method. As an example, the power and feasibility of the method are assessed on simulated data for a simplified version of an off-shell Higgs boson couplings measurement in the four-lepton final states. This approach represents an extension to the standard statistical methodology used by the experiments at the Large Hadron Collider, and can benefit many physics analyses.

frequentist statistics

Measurement of off-shell Higgs boson production in the $H^*\rightarrow ZZ\rightarrow 4\ell$ decay channel using a neural simulation-based inference technique in 13 TeV pp collisions with the ATLAS detector

A measurement of off-shell Higgs boson production in the $H^*\to ZZ\to 4\ell$ decay channel is presented. The measurement uses 140 fb −1 of proton–proton collisions at $\sqrt{s} = 13$ TeV collected by the ATLAS detector at the Large Hadron Collider and supersedes the previous result in this decay channel using the same dataset. The data analysis is performed using a neural simulation-based inference method, which builds per-event likelihood ratios using neural networks. The observed (expected) off-shell Higgs boson production signal strength in the $ZZ\to 4\ell$ decay channel at 68% CL is $0.87^{+0.75}_{-0.54}$ ($1.00^{+1.04}_{-0.95}$ ). The evidence for off-shell Higgs boson production using the $ZZ\to 4\ell$ decay channel has an observed (expected) significance of 2.5σ (1.3σ). The expected result represents a significant improvement relative to that of the previous analysis of the same dataset, which obtained an expected significance of 0.5σ. When combined with the most recent ATLAS measurement in the $ZZ\to 2\ell 2\nu$ decay channel, the evidence for off-shell Higgs boson production has an observed (expected) significance of 3.7σ (2.4σ). The off-shell measurements are combined with the measurement of on-shell Higgs boson production to obtain constraints on the Higgs boson total width. The observed (expected) value of the Higgs boson width at 68% CL is $4.3^{+2.7}_{-1.9}$ ($4.1^{+3.5}_{-3.4}$ ) MeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Cosmological neutrino mass: a frequentist overview in light of DESI

We derive constraints on the neutrino mass using a variety of recent cosmological datasets, including DESI BAO, the full-shape analysis of the DESI matter power spectrum and the one-dimensional power spectrum of the Lyman-α forest (P1D) from eBOSS quasars as well as the cosmic microwave background (CMB). The constraints are obtained in the frequentist formalism by constructing profile likelihoods and applying the Feldman-Cousins prescription to compute confidence intervals. This method avoids potential prior and volume effects that may arise in a comparable Bayesian analysis. Parabolic fits to the profiles allow one to distinguish changes in the upper limits from variations in the constraining power σ of the different data combinations. We find that all profiles in the ΛCDM model are cut off by the ∑m ν ≥ 0 bound, meaning that the corresponding parabolas reach their minimum in the unphysical sector. The most stringent 95% C.L. upper limit is obtained by the combination of DESI DR2 BAO, Planck PR4 and CMB lensing at 53 meV, below the minimum of 59 meV set by the normal ordering. The corresponding constraining power σ is 43 meV, which highlights the importance of the cut-off by negative values in the determination of the upper limit. Extending ΛCDM to non-zero curvature and w 0 w a CDM relaxes the constraints past 59 meV again, but only w 0 w a CDM exhibits profiles with a minimum at a positive value. Additionally, we extend the formalism to constrain the lightest neutrino mass. For DESI DR2 BAO, Planck PR4 and CMB lensing, we find confidence limits at 20 and 19 meV for normal and inverted ordering, respectively. Using a combination of DESI DR1 full-shape, BBN and eBOSS Lyman-α P1D, we successfully constrain the neutrino mass independently of the CMB. This combination yields m l ≤ 97 and 98 meV in the normal and inverted orderings, and total neutrino mass ∑m ν ≤ 285 meV (95% C.L.). The addition of DESI full-shape or Lyman-α P1D to CMB and DESI BAO results in small but noticeable improvement of the constraining power of the data. Lyman-α free-streaming measurements especially improve the constraint. Since they are based on eBOSS data, this sets a promising precedent for upcoming DESI data.

Frequentist statistics

Frequentist cosmological constraints from full-shape clustering measurements in DESI DR1

We present a frequentist analysis of clustering measurements from Data Release 1 of the Dark Energy Spectroscopic Instrument (DESI) using the standard profile likelihood method. While Bayesian inferences for effective field theory models of galaxy clustering can be highly sensitive to prior choices for extended cosmological models, frequentist inferences are not susceptible to such effects. We compare frequentist and Bayesian constraints for the parameter set {σ 8 , H 0 , Ω m , w 0 , w a } using the full-shape power spectrum multipoles, post-reconstruction baryon acoustic oscillation (BAO) measurements, and external datasets from the CMB and type Ia supernovae measurements. The frequentist confidence intervals are significantly shifted relative to the Bayesian credible intervals for the w 0 w a CDM model, unless supernovae data are included. When DESI full-shape and BAO data are fit jointly, we obtain the following 1σ frequentist confidence intervals for ΛCDM (w 0 w a CDM): σ 8 = 0.863 +0.048 -0.040 , H 0 = 68.96 +0.81 -0.80 km s -1 Mpc -1 , Ω m = 0.3034 ± 0.0110 (σ 8 = 0.782 +0.060 -0.036 , H 0 = 63.7 +4.2 -2.0 km s -1 Mpc -1 , Ω m = 0.378 +0.024 -0.047 , w 0 = -0.16 +0.10 -0.50 , w a = -3.0 +1.7 ), corresponding to 0.8σ, 0.3σ, 0.7σ (2.1σ, 4.1σ, 6.5σ, 6.3σ, 6.6σ) shifts between the maximum likelihood estimate and the Bayesian posterior mean for ΛCDM (w 0 w a CDM) respectively.

Bayesian reasoning

Improvement of Drop‐Hammer Impact Testing for Safety Assessment of High Explosives Using 10‐mg Samples

Here, in this study, we established an improved method for drop-hammer impact testing of small quantities of high explosives (10 mg). We performed about seven hundred impact tests under various experimental conditions (e.g., sandpaper vs bare anvil, different sample masses, drop-weights, and striker diameters) to determine an optimal set of conditions and reaction detection methods (e.g., gas analysis, video, and sound recordings) that give the most statistically reliable results with 10 mg samples. We used both Frequentist and Bayesian statistical approaches to compare estimates of the drop height (DH50) that initiates a reaction 50% of the time, and to quantify the associated uncertainty. Gas analysis proved to be the most reliable reaction detection method, showing unambiguous rises in HE decomposition products (e.g., CO 2 ) even when the other indicators (e.g., sound, video) were inconclusive. The impact tests performed with a bare anvil showed much better reproducibility than those conducted with sandpaper, reducing the largest uncertainty observed in the data sets by a factor of 1.7. The DH 50 values obtained from three different sample masses (10, 20, and 35 mg) fell within the uncertainties of the measurements. We demonstrated the improved procedure (i.e., 10-mg samples, gas analysis, bare anvil, and Bayesian approach) on a variety of PETN samples having different surface areas and thermal histories.

PETN

A novel framework for increasing research transparency: Exploring the connection between diversity and innovation

A split sample/dual method research protocol is demonstrated to increase transparency while reducing the probability of false discovery. We apply the protocol to examine whether diversity in ownership teams increases or decreases the likelihood of a firm reporting a novel innovation using data from the 2018 United States Census Bureau’s Annual Business Survey. Transparency is increased in three ways: 1) all specification testing and identifying potentially productive models is done in an exploratory subsample that 2) preserves the validity of hypothesis test statistics fromde novoestimation in the holdout confirmatory sample with 3) all findings publicly documented in an earlier registered report and in this journal publication. Bayesian estimation procedures that leverage information from the exploratory stage included in the confirmatory stage estimation replace traditional frequentist null hypothesis significance testing. In addition to increasing statistical power by using information from the full sample, Bayesian methods directly estimate a probability distribution for the magnitude of an effect, allowing much richer inference. Estimated magnitudes of diversity along academic discipline, race, ethnicity, and foreign-born status dimensions are positively associated with innovation. A maximally diverse ownership team on these dimensions would be roughly six times more likely to report new-to-market innovation than a homophilic team.

Science & Technology - Other Topics

Role of the likelihood for elastic scattering uncertainty quantification

In the last decade, uncertainty quantification (UQ) for optical model potentials (OMPs) has become a focal point for nuclear reaction theory, and several competing approaches for OMP UQ have recently been developed. Here, we clarify recent efforts to compare frequentist and Bayesian approaches in the context of OMP UQ [G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019)]. We replicate a portion of that OMP UQ study but use independent statistical tools. Specifically, we compare two methods for OMP parameter inference from elastic scattering data: the Levenberg-Marquardt algorithm for χ 2 minimization on one hand and Markov chain Monte Carlo (MCMC) sampling on the other. Separately, we assess the common practice of using a renormalized likelihood (χ 2 /N), N being the number of data points, instead of the canonical weighted-least-squares likelihood (χ 2 ), as a way of accounting for unknown data correlations. Here, we show that for a generic linear model and for a five-parameter OMP analysis, frequentist and uniform-prior Bayesian approaches recover the same optimum and uncertainty estimates—not systematically larger uncertainties for the Bayesian approach, as was concluded in G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019). Further, we show that if an additional, near-degenerate parameter is introduced into the same OMP analysis such that the parameter posterior becomes non-Gaussian, then covariance-based estimates of uncertainty become unreliable. Finally, we show that regardless of optimization approach, if χ 2 /N is used for the likelihood, the resulting parametric uncertainties increase by $\sqrt{N}$, and that this is responsible for the conclusions drawn in the revisited study. Based on our replication results, we find that a fortuitous cancellation of unreported errors and the renormalization factor can lead to improvement in empirical coverages, as was the case in the original comparative study. We emphasize that developing and applying a realistic likelihood function is an essential task in a UQ analysis, and that several recent UQ studies that employed a renormalized likelihood (i.e., including a 1/N factor) may have yielded unrealistically large uncertainties for elastic-scattering observables. If the parameter posterior deviates from multivariate-normal, a sampling-based approach like MCMC has a clear advantage over methods that assume the Laplace approximation holds. We note that empirical coverage can serve as an important internal check for the analyst whose model or data may have additional, unaccounted-for uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Alleviating prior dependencies for DESI DR1 clustering fits through reparameterization

Bayesian analyses of the full-shape clustering of Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) exhibit prior-volume projection effects, whereby weakly constrained nuisance parameters of the Effective Field Theory of Large Scale Structure (EFTofLSS) shift marginalized cosmological posteriors away from the posterior maximum. We reanalyze DESI DR1 power spectrum multipoles using two complementary mitigation strategies: (i) nonlinear orthogonalization to decorrelate nuisance and cosmological parameter priors, and (ii) a fully reparameterization-invariant Jeffreys prior over all EFTofLSS coefficients, evaluated on-the-fly via closed-form Jacobians. Including data from DESI, Big-Bang Nuclesynthesis and a constraint on $n_{\mathrm{s}}$, baseline priors lead to multi-$σ$ projection in the Hubble parameter $H_{0}$ and dark energy equation of state parameters $w_{0}$ and $w_{a}$; the Jeffreys prior successfully recenters these posteriors to enclose the maximum a posteriori estimate within the 68% credible regions, demonstrating clear mitigation of projection effects for these late-time expansion parameters. A hybrid Jeffreys+baseline-Gaussian configuration controls residual over-broad tails in the physical cold dark matter density $ω_{\mathrm{c}}$ while preserving the volume correction, and is our favoured approach. We compare the credible intervals derived using our methodology to those obtained using Halo Occupation Distribution (HOD)-informed priors and to confidence intervals derived using frequentist profile likelihood analyses, finding agreement in both central values and degeneracy directions in the $w_{0}$--$w_{a}$ plane. This demonstrates that, once projection effects are properly controlled, we can make robust inferences about the late-time cosmological expansion independent of the statistical framework adopted.

Bonici, M. [Waterloo U.; Perimeter Inst. Theor. Ph

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip

Informed total-error-minimizing priors: Interpretable cosmological parameter constraints despite complex nuisance effects

While Bayesian inference techniques are standard in cosmological analyses, it is common to interpret resulting parameter constraints with a frequentist intuition. This intuition can fail, for example, when marginalizing high-dimensional parameter spaces onto subsets of parameters, because of what has come to be known as projection effects or prior volume effects. We present the method of informed total-error-minimizing (ITEM) priors to address this problem. An ITEM prior is a prior distribution on a set of nuisance parameters, such as those describing astrophysical or calibration systematics, intended to enforce the validity of a frequentist interpretation of the posterior constraints derived for a set of target parameters (e.g., cosmological parameters). Our method works as follows. For a set of plausible nuisance realizations, we generate target parameter posteriors using several different candidate priors for the nuisance parameters. We reject candidate priors that do not accomplish the minimum requirements of bias (of point estimates) and coverage (of confidence regions among a set of noisy realizations of the data) for the target parameters on one or more of the plausible nuisance realizations. Of the priors that survive this cut, we select the ITEM prior as the one that minimizes the total error of the marginalized posteriors of the target parameters. As a proof of concept, we applied our method to the density split statistics measured in Dark Energy Survey Year 1 data. We demonstrate that the ITEM priors substantially reduce prior volume effects that otherwise arise and that they allow for sharpened yet robust constraints on the parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm