Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “likelihood”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Mixture density network estimation of continuous variable maximum likelihood using discrete training samples

Abstract Mixture density networks (MDNs) can be used to generate posterior density functions of model parameters $$\varvec{\theta }$$ θ given a set of observables $${\mathbf {x}}$$ x . In some applications, training data are available only for discrete values of a continuous parameter $$\varvec{\theta }$$ θ . In such situations, a number of performance-limiting issues arise which can result in biased estimates. We demonstrate the usage of MDNs for parameter estimation, discuss the origins of the biases, and propose a corrective method for each issue.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Development of systematic uncertainty-aware neural network trainings for binned-likelihood analyses at the LHC

We propose a neural network training method capable of accounting for the effects of systematic variations of the data model in the training process and describe its extension towards neural network multiclass classification. The procedure is evaluated on the realistic case of the measurement of Higgs boson production via gluon fusion and vector boson fusion in the τ τ decay channel at the CMS experiment. The neural network output functions are used to infer the signal strengths for inclusive production of Higgs bosons as well as for their production via gluon fusion and vector boson fusion. We observe improvements of 12 and 16% in the uncertainty in the signal strengths for gluon and vector-boson fusion, respectively, compared with a conventional neural network training based on cross-entropy.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Beyond Fisher forecasting for cosmology

The planning and design of future experiments rely heavily on forecasting to assess the potential scientific value provided by a hypothetical set of measurements. The Fisher information matrix, due to its convenient properties and low computational cost, provides an especially useful forecasting tool. However, the Fisher matrix only provides a reasonable approximation to the true likelihood when data are nearly Gaussian distributed and observables have nearly linear dependence on the parameters of interest. Also, Fisher forecasting techniques alone cannot be used to assess their own validity. Thorough sampling of the exact or mock likelihood can definitively determine whether a Fisher forecast is valid, though such sampling is often prohibitively expensive. Here we propose a simple test, based on the Derivative Approximation for likelihoods (DALI) technique, to determine whether the Fisher matrix provides a good approximation to the exact likelihood. We show that the Fisher matrix becomes a poor approximation to the true likelihood in regions where two-dimensional slices of level surfaces of the DALI approximation to the likelihood differ from two-dimensional slices of level surfaces of the Fisher approximation to the likelihood. We demonstrate that our method accurately predicts situations in which the Fisher approximation deviates from the true likelihood for various cosmological models and several data combinations, with only a modest increase in computational cost compared to standard Fisher forecasts.

79 ASTRONOMY AND ASTROPHYSICS↗

Multiscale Flow for robust and optimal cosmological analysis

We propose Multiscale Flow, a generative Normalizing Flow that creates samples and models the field-level likelihood of two-dimensional cosmological data such as weak lensing. Multiscale Flow uses hierarchical decomposition of cosmological fields via a wavelet basis and then models different wavelet components separately as Normalizing Flows. The log-likelihood of the original cosmological field can be recovered by summing over the log-likelihood of each wavelet term. This decomposition allows us to separate the information from different scales and identify distribution shifts in the data such as unknown scale-dependent systematics. The resulting likelihood analysis can not only identify these types of systematics, but can also be made optimal, in the sense that the Multiscale Flow can learn the full likelihood at the field without any dimensionality reduction. We apply Multiscale Flow to weak lensing mock datasets for cosmological inference and show that it significantly outperforms traditional summary statistics such as power spectrum and peak counts, as well as machine learning–based summary statistics such as scattering transform and convolutional neural networks. We further show that Multiscale Flow is able to identify distribution shifts not in the training data such as baryonic effects. Finally, we demonstrate that Multiscale Flow can be used to generate realistic samples of weak lensing data.

79 ASTRONOMY AND ASTROPHYSICS↗

Translation and rotation equivariant normalizing flow (TRENF) for optimal cosmological analysis

ABSTRACT Our Universe is homogeneous and isotropic, and its perturbations obey translation and rotation symmetry. In this work, we develop translation and rotation equivariant normalizing flow (TRENF), a generative normalizing flow (NF) model which explicitly incorporates these symmetries, defining the data likelihood via a sequence of Fourier space-based convolutions and pixel-wise non-linear transforms. TRENF gives direct access to the high dimensional data likelihood p(x|y) as a function of the labels y, such as cosmological parameters. In contrast to traditional analyses based on summary statistics, the NF approach has no loss of information since it preserves the full dimensionality of the data. On Gaussian random fields, the TRENF likelihood agrees well with the analytical expression and saturates the Fisher information content in the labels y. On non-linear cosmological overdensity fields from N-body simulations, TRENF leads to significant improvements in constraining power over the standard power spectrum summary statistic. TRENF is also a generative model of the data, and we show that TRENF samples agree well with the N-body simulations it trained on, and that the inverse mapping of the data agrees well with a Gaussian white noise both visually and on various summary statistics: when this is perfectly achieved the resulting p(x|y) likelihood analysis becomes optimal. Finally, we develop a generalization of this model that can handle effects that break the symmetry of the data, such as the survey mask, which enables likelihood analysis on data without periodic boundaries.

79 ASTRONOMY AND ASTROPHYSICS↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

Adaptive hyperparameter updating for training restricted Boltzmann machines on quantum annealers

Restricted Boltzmann Machines (RBMs) have been proposed for developing neural networks for a variety of unsupervised machine learning applications such as image recognition, drug discovery, and materials design. The Boltzmann probability distribution is used as a model to identify network parameters by optimizing the likelihood of predicting an output given hidden states trained on available data. Training such networks often requires sampling over a large probability space that must be approximated during gradient based optimization. Quantum annealing has been proposed as a means to search this space more efficiently which has been experimentally investigated on D-Wave hardware. D-Wave implementation requires selection of an effective inverse temperature or hyperparameter (β) within the Boltzmann distribution which can strongly influence optimization. Here, we show how this parameter can be estimated as a hyperparameter applied to D-Wave hardware during neural network training by maximizing the likelihood or minimizing the Shannon entropy. We find both methods improve training RBMs based upon D-Wave hardware experimental validation on an image recognition problem. Neural network image reconstruction errors are evaluated using Bayesian uncertainty analysis which illustrate more than an order magnitude lower image reconstruction error using the maximum likelihood over manually optimizing the hyperparameter. The maximum likelihood method is also shown to out-perform minimizing the Shannon entropy for image reconstruction.

97 MATHEMATICS AND COMPUTING↗

Statistical interpretation of sterile neutrino oscillation searches at reactors

A considerable experimental effort is currently under way to test the persistent hints for oscillations due to an eV-scale sterile neutrino in the data of various reactor neutrino experiments. The assessment of the statistical significance of these hints is usually based on Wilks’ theorem, whereby the assumption is made that the log-likelihood is χ 2 -distributed. However, it is well known that the preconditions for the validity of Wilks’ theorem are not fulfilled for neutrino oscillation experiments. In this work we derive a simple asymptotic form of the actual distribution of the log-likelihood based on reinterpreting the problem as fitting white Gaussian noise. From this formalism we show that, even in the absence of a sterile neutrino, the expectation value for the maximum likelihood estimate of the mixing angle remains non-zero with attendant large values of the log-likelihood. Our analytical results are then confirmed by numerical simulations of a toy reactor experiment. Finally, we apply this framework to the data of the Neutrino-4 experiment and show that the null hypothesis of no-oscillation is rejected at the 2.6 σ level, compared to 3.2 σ obtained under the assumption that Wilks’ theorem applies.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Legacy Analysis of Dark Matter Annihilation from the Milky Way Dwarf Spheroidal Galaxies with 14 Years of Fermi-LAT Data: Data Products

https://arxiv.org/abs/2311.04982This repository contains data products from the 14 year Fermi-LAT analysis of the Milky Way dSphs as described in McDaniel et al (2023) https://arxiv.org/abs/2311.04982. The included data products are the SED fits files and 2D TS profiles in the WIMP mass and cross section space. These are available for dSphs as well as the blank-field analysis, using both the standard likelihood and the weighted likelihood approach (see Appendix A). CSV data files are included containing relevant information about the dSphs (see table 1 of McDaniel+2024) and blank fields (RA & Dec). Also included is a python Jupyter Notebook to show basic usage of the data products. For Example, plotting the SED likelihoods, converting the SED likelihood to DM space, creating a combined TS profile, obtaining upper limits, etc. This is also stored as a static html file for easier viewing. SEDs are stored as fits files in the output format of fermipy (see https://fermipy.readthedocs.io/en/latest/advanced/sed.html) TS profiles are stored as numpy arrays covering 40 logarithmically spaced mass values over the 1 GeV - 1 TeV mass range and 60 logarithmically spaced cross-section values covering the $10^{-28}$ to $10^{-22}$ cm$^3$/s cross section range. TS profiles including the J-factor prior and without are available, and are labeled with "Jprior" or "noprior" respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Self-Supervised Anomaly Detection via Neural Autoregressive Flows with Active Learning

Many self-supervised methods have been proposed with the target of image anomaly detection. These methods often rely on the paradigm of data augmentation with predefined transformations such as flipping, cropping, and rotations. However, it is not straightforward to apply these techniques for non-image data, such as time series or tabular data, while the performance of the existing deep approaches has been under our expectation on tasks beyond images. In this work, we propose a novel active learning (AL) scheme that relied on neural autoregressive flows (NAF) for self-supervised anomaly detection, specifically on small-scale data. Unlike other generative models such as GANs or VAEs, flow-based models allow to explicitly learn the probability density and thus can assign accurate likelihoods to normal data which makes it usable to detect anomalies. The proposed NAF-AL method is achieved by efficiently generating random samples from latent space and transforming them into feature space along with likelihoods via invertible mapping. The samples with lower likelihoods are selected and further checked by outlier detection using Mahalanobis distance. The augmented samples incorporating with normal samples are used for training a better detector so as to approach decision boundaries. Compared with random transformations, NAF-AL can be interpreted as a likelihood-oriented data augmentation that is more efficient and robust. Extensive experiments show that our approach outperforms existing baselines on multiple time series and tabular datasets, and a real-world application in advanced manufacturing, with significant improvement on anomaly detection accuracy and robustness over the state-of-the-art.

Zhang, Jiaxin↗

A Primer on Dose-Response Data Modeling in Radiation Therapy

An overview of common approaches used to assess a dose response for radiation therapy–associated endpoints is presented, using lung toxicity data sets analyzed as a part of the High Dose per Fraction, Hypofractionated Treatment Effects in the Clinic effort as an example. Each component presented (eg, data-driven analysis, dose-response analysis, and calculating uncertainties on model prediction) is addressed using established approaches. Specifically, the maximum likelihood method was used to calculate best parameter values of the commonly used logistic model, the profile-likelihood to calculate confidence intervals on model parameters, and the likelihood ratio to determine whether the observed data fit is statistically significant. The bootstrap method was used to calculate confidence intervals for model predictions. Correlated behavior of model parameters and implication for interpreting dose response are discussed.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Fast matrix algebra for Bayesian model calibration

In Bayesian model calibration, evaluation of the likelihood function usually involves finding the inverse and determinant of a covariance matrix. When Markov Chain Monte Carlo (MCMC) methods are used to sample from the posterior, hundreds of thousands of likelihood evaluations may be required. In this paper, we demonstrate that the structure of the covariance matrix can be exploited, leading to substantial time savings in practice. Here, we also derive two simple equations for approximating the inverse of the covariance matrix in this setting, which can be computed in near-quadratic time. The practical implications of these strategies are demonstrated using a simple numerical case study and the "quack" R package. For a covariance matrix with 1000 rows, application of these strategies for a million likelihood evaluations leads to a speedup of roughly 4000 compared to the naive implementation

97 MATHEMATICS AND COMPUTING↗