Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “non-Gaussian error distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Computational Information Criterion for Particle-Tracking with Sparse or Noisy Data

Traditional probabilistic methods for the simulation of advection-diffusion equations (ADEs) often overlook the entropic contribution of the discretization, e.g., the number of particles, within associated numerical methods. Many times, the gain in accuracy of a highly discretized numerical model is outweighed by its associated computational costs or the noise within the data. Herein, we address the question of how many particles are needed in a simulation to best approximate and estimate parameters in one-dimensional advective-diffusive transport. To do so, we use the well-known Akaike Information Criterion (AIC) and a recently-developed correction called the Computational Information Criterion (COMIC) to guide the model selection process. Random-walk and mass-transfer particle tracking methods are employed to solve the model equations at various levels of discretization. Numerical results demonstrate that the COMIC provides an optimal number of particles that can describe a more efficient model in terms of parameter estimation and model prediction compared to the model selected by the AIC even when the data is sparse or noisy, the sampling volume is not uniform throughout the physical domain, or the error distribution of the data is non-IID Gaussian.

97 MATHEMATICS AND COMPUTING↗

Forecasting constraints on the mean free path of ionizing photons at z ≥ 5.4 from the Lyman-α forest flux autocorrelation function

ABSTRACT Fluctuations in Lyman-α (Ly α) forest transmission towards high-z quasars are partially sourced from spatial fluctuations in the ultraviolet background, the level of which are set by the mean free path of ionizing photons (λmfp). The autocorrelation function of Ly α forest flux characterizes the strength and scale of transmission fluctuations and, as we show, is thus sensitive to λmfp. Recent measurements at z ∼ 6 suggest a rapid evolution of λmfp at z > 5.0 which would leave a signature in the evolution of the autocorrelation function. For this forecast, we model mock Ly α forest data with properties similar to the XQR-30 extended data set at 5.4 ≤ z ≤ 6.0. At each z, we investigate 100 mock data sets and an ideal case where mock data matches model values of the autocorrelation function. For ideal data with λmfp = 9.0 cMpc at z = 6.0, we recover $\lambda _{\text{mfp}}=12^{+6}_{-3}$ cMpc. This precision is comparable to direct measurements of λmfp from the stacking of quasar spectra beyond the Lyman limit. Hypothetical high-resolution data leads to a $\sim 40~{{\ \rm per\ cent}}$ reduction in the error bars over all z. The distribution of mock values of the autocorrelation function in this work is highly non-Gaussian for high-z, which should caution work with other statistics of the high-z Ly α forest against making this assumption. We use a rigorous statistical method to pass an inference test, however future work on non-Gaussian methods will enable higher precision measurements.

Astronomy & Astrophysics↗

Sequential Kalman tuning of the t -preconditioned Crank-Nicolson algorithm: efficient, adaptive and gradient-free inference for Bayesian inverse problems

Ensemble Kalman Inversion (EKI) has been proposed as an efficient method for the approximate solution of Bayesian inverse problems with expensive forward models. However, when applied to the Bayesian inverse problem EKI is only exact in the regime of Gaussian target measures and linear forward models. Here, in this work we propose embedding EKI and Flow Annealed Kalman Inversion, its normalizing flow (NF) preconditioned variant, within a Bayesian annealing scheme as part of an adaptive implementation of the t-preconditioned Crank-Nicolson (tpCN) sampler. The tpCN sampler differs from standard pCN in that its proposal is reversible with respect to the multivariate t-distribution. The more flexible tail behaviour allows for better adaptation to sampling from non-Gaussian targets. Within our Sequential Kalman Tuning (SKT) adaptation scheme, EKI is used to initialize and precondition the tpCN sampler for each annealed target. The subsequent tpCN iterations ensure particles are correctly distributed according to each annealed target, avoiding the accumulation of errors that would otherwise impact EKI. We demonstrate the performance of SKT for tpCN on three challenging numerical benchmarks, showing significant improvements in the rate of convergence compared to adaptation within standard SMC with importance weighted resampling at each temperature level, and compared to similar adaptive implementations of standard pCN. The SKT scheme applied to tpCN offers an efficient, practical solution for solving the Bayesian inverse problem when gradients of the forward model are not available. Code implementing the SKT schemes for tpCN is available at https://github.com/RichardGrumitt/KalmanMC.

97 MATHEMATICS AND COMPUTING↗

Type Ia Supernova Growth-rate Measurement with LSST Simulations: Intrinsic Scatter Systematics

Measurement of the growth rate of structures (fσ 8 ) with Type Ia supernovae (SNe Ia) will improve our understanding of the nature of dark energy and enable tests of general relativity. In this paper, we generate simulations of the 10 yr SN Ia data set of the Rubin-LSST survey, including a correlated velocity field from an N-body simulation and realistic models of SNe Ia properties and their correlations with host-galaxy properties. We find, similar to SN Ia analyses that constrain the dark energy equation-of-state parameters w 0 w a , that constraints on fσ 8 can be biased depending on the intrinsic scatter of SNe Ia. While for the majority of intrinsic scatter models we recover fσ 8 with a precision of ∼13%–14%, for the most realistic dust-based model, we find that the presence of non-Gaussianities in Hubble diagram residuals leads to a bias on fσ 8 of ∼ −20%. When trying to correct for the dust-based intrinsic scatter, we find that the propagation of the uncertainty on the model parameters does not significantly increase the error on fσ 8 . We also find that while the main component of the error budget of fσ 8 is the statistical uncertainty (>75% of the total error budget), the systematic error budget is dominated by the uncertainty on the damping parameter, σ u , that gives an empirical description of the effect of redshift space distortions on the velocity power spectrum. Our results motivate a search for new methods to correct for the non-Gaussian distribution of the Hubble diagram residuals, as well as an improved modeling of the damping parameter.

Carreres, Bastien [Duke Univ., Durham, NC (United ↗

BayFlux: A Bayesian method to quantify metabolic Fluxes and their uncertainty at the genome scale

Metabolic fluxes, the number of metabolites traversing each biochemical reaction in a cell per unit time, are crucial for assessing and understanding cell function. 13 C Metabolic Flux Analysis ( 13 C MFA) is considered to be the gold standard for measuring metabolic fluxes. 13 C MFA typically works by leveraging extracellular exchange fluxes as well as data from 13 C labeling experiments to calculate the flux profile which best fit the data for a small, central carbon, metabolic model. However, the nonlinear nature of the 13 C MFA fitting procedure means that several flux profiles fit the experimental data within the experimental error, and traditional optimization methods offer only a partial or skewed picture, especially in “non-gaussian” situations where multiple very distinct flux regions fit the data equally well. Here, we present a method for flux space sampling through Bayesian inference (BayFlux), that identifies the full distribution of fluxes compatible with experimental data for a comprehensive genome-scale model. This Bayesian approach allows us to accurately quantify uncertainty in calculated fluxes. We also find that, surprisingly, the genome-scale model of metabolism produces narrower flux distributions (reduced uncertainty) than the small core metabolic models traditionally used in 13 C MFA. The different results for some reactions when using genome-scale models vs core metabolic models advise caution in assuming strong inferences from 13 C MFA since the results may depend significantly on the completeness of the model used. Based on BayFlux, we developed and evaluated novel methods (P- 13 C MOMA and P- 13 C ROOM) to predict the biological results of a gene knockout, that improve on the traditional MOMA and ROOM methods by quantifying prediction uncertainty.

59 BASIC BIOLOGICAL SCIENCES↗

Earth System Reanalysis in Support of Climate Model Improvements

Recent climate model developments, established through increased model resolution, have led to substantial improvements in model simulations of the time-evolving, coupled Earth system and its subcomponents. However, regardless of resolution, climate models will always produce climate features and variability that differ from the real world and will be prone to biases. This is due to many remaining uncertainties, such as in parametric and structural model uncertainty, in the initial conditions prescribed, and in the prescribed (scenario) forcing which varies on decadal to centennial timescales. Further model improvements are expected to arise specifically from improved representation of physical processes realized through model-data fusion. This will create an unprecedented opportunity to better exploit a large array of Earth observations, from in situ measurements to weather radars and satellite observations, as the resolved scales of the models approach those of the observations. For this, climate DA will be the central tool to bring models and observations into consistency, by improving initial conditions, inferring uncertain model parameters and structure, and quantifying uncertainty. Generally, there will be advantages and complementarities of adjoint-based smoother approaches, ensemble-based filter approaches, or new ML-inspired approaches. Yet, the ever-increasing model resolution will present growing challenges arising from computational cost, calling for new ways of performing data assimilation and model optimization. Using the complementarity in a hybrid approach, blending tools and concepts from variational, ensemble and ML methods might be what is required in the future. In this context ML could be important to handle non-linear responses, and to better approximate non-Gaussian distributions.

54 ENVIRONMENTAL SCIENCES↗

Primordial non-Gaussianity from the completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey – I: Catalogue preparation and systematic mitigation

ABSTRACT We investigate the large-scale clustering of the final spectroscopic sample of quasars from the recently completed extended Baryon Oscillation Spectroscopic Survey (eBOSS). The sample contains 343 708 objects in the redshift range 0.8 < z < 2.2 and 72 667 objects with redshifts 2.2 < z < 3.5, covering an effective area of $4699\, {\rm deg}^{2}$. We develop a neural network-based approach to mitigate spurious fluctuations in the density field caused by spatial variations in the quality of the imaging data used to select targets for follow-up spectroscopy. Simulations are used with the same angular and radial distributions as the real data to estimate covariance matrices, perform error analyses, and assess residual systematic uncertainties. We measure the mean density contrast and cross-correlations of the eBOSS quasars against maps of potential sources of imaging systematics to address algorithm effectiveness, finding that the neural network-based approach outperforms standard linear regression. Stellar density is one of the most important sources of spurious fluctuations, and a new template constructed using data from the Gaia spacecraft provides the best match to the observed quasar clustering. The end-product from this work is a new value-added quasar catalogue with the improved weights to correct for non-linear imaging systematic effects, which will be made public. Our quasar catalogue is used to measure the local-type primordial non-Gaussianity in a companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Flow-based likelihoods for non-Gaussian inference

We investigate the use of data-driven likelihoods to bypass a key assumption made in many scientific analyses, which is that the true likelihood of the data is Gaussian. In particular, we suggest using the optimization targets of flow-based generative models, a class of models that can capture complex distributions by transforming a simple base distribution through layers of nonlinearities. We call these flow-based likelihoods (FBL). We analyze the accuracy and precision of the reconstructed likelihoods on mock Gaussian data, and show that simply gauging the quality of samples drawn from the trained model is not a sufficient indicator that the true likelihood has been learned. We nevertheless demonstrate that the likelihood can be reconstructed to a precision equal to that of sampling error due to a finite sample size. We then apply FBLs to mock weak lensing convergence power spectra, a cosmological observable that is significantly non-Gaussian (NG). We find that the FBL captures the NG signatures in the data extremely well, while other commonly used data-driven likelihoods, such as Gaussian mixture models and independent component analysis, fail to do so. This suggests that works that have found small posterior shifts in NG data with data-driven likelihoods such as these could be underestimating the impact of non-Gaussianity in parameter constraints. By introducing a suite of tests that can capture different levels of NG in the data, we show that the success or failure of traditional data-driven likelihoods can be tied back to the structure of the NG in the data. Here, unlike other methods, the flexibility of the FBL makes it successful at tackling different types of NG simultaneously. Because of this, and consequently their likely applicability across datasets and domains, we encourage their use for inference when sufficient mock data are available for training.

79 ASTRONOMY AND ASTROPHYSICS↗

Delensing the CMB with the cosmic infrared background: the impact of foregrounds

ABSTRACT The most promising avenue for detecting primordial gravitational waves from cosmic inflation is through measurements of degree-scale cosmic microwave background (CMB) B-mode polarization. This approach must face the challenge posed by gravitational lensing of the CMB, which obscures the signal of interest. Fortunately, the lensing effects can be partially removed by combining high-resolution E-mode measurements with an estimate of the projected matter distribution. For near-future experiments, the best estimate of the latter will arise from co-adding internal reconstructions (derived from the CMB itself) with external tracers such as the cosmic infrared background (CIB). In this work, we characterize how foregrounds impact the delensing procedure when CIB intensity, I, is used as the matter tracer. We find that higher point functions of the CIB and Galactic dust such as 〈BEI〉c and 〈EIEI〉c can, in principle, bias the power spectrum of delensed B-modes. To quantify these, we first estimate the dust residuals in currently available CIB maps and upcoming, foreground-cleaned Simons Observatory CMB data. Then, using non-Gaussian simulations of Galactic dust – extrapolated to the relevant frequencies, assuming the spectral index of polarized dust emission to be fixed at the value determined by Planck – we show that the bias to any primordial signal is small compared to statistical errors for ground-based experiments, but might be significant for space-based experiments probing very large angular scales. However, mitigation techniques based on multifrequency cleaning appear to be very effective. We also show, by means of an analytical model, that the bias arising from the higher point functions of the CIB itself ought to be negligible.

79 ASTRONOMY AND ASTROPHYSICS↗

Suppressing the sample variance of DESI-like galaxy clustering with fast simulations

Ongoing and upcoming galaxy redshift surveys, such as the Dark Energy Spectroscopic Instrument (DESI) survey, will observe vast regions of sky and a wide range of redshifts. In order to model the observations and address various systematic uncertainties, N-body simulations are routinely adopted, however, the number of large simulations with sufficiently high mass resolution is usually limited by available computing time. Therefore, achieving a simulation volume with the effective statistical errors significantly smaller than those of the observations becomes prohibitively expensive. In this study, we apply the Convergence Acceleration by Regression and Pooling (CARPool) method to mitigate the sample variance of the DESI-like galaxy clustering in the AbacusSummit simulations, with the assistance of the quasi-N-body simulations FastPM. Based on the halo occupation distribution (HOD) models, we construct different FastPM galaxy catalogs, including the luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars, with their number densities and two-point clustering statistics well matched to those of AbacusSummit. We also employ the same initial conditions between AbacusSummit and FastPM to achieve high cross-correlation, as it is useful in effectively suppressing the variance. Our method of reducing noise in clustering is equivalent to performing a simulation with volume larger by a factor of 5 and 4 for LRGs and ELGs, respectively. We also mitigate the standard deviation of the LRG bispectrum with the triangular configurations k 2 = 2k 1 = 0.2 h Mpc -1 by a factor of 1.6. With smaller sample variance on galaxy clustering, we are able to constrain the baryon acoustic oscillations (BAO) scale parameters to higher precision. The CARPool method will be beneficial to better constrain the theoretical systematics of BAO, redshift space distortions (RSD) and primordial non-Gaussianity (NG).

79 ASTRONOMY AND ASTROPHYSICS↗

Globally optimal interferometry with lossy twin Fock probes

Parity or quadratic spin (e.g., J z 2 ) readouts of a Mach–Zehnder (MZ) interferometer probed with a twin Fock (TF) input state allow saturating the optimal sensitivity attainable among all mode-separable states with a fixed total number of particles but only when the interferometer phase θ is near zero. When more general Dicke state probes are used, the parity readout saturates the quantum Fisher information (QFI) at θ = 0, whereas better-than-standard quantum limit performance of the J z 2 readout is restricted to an o ( N ) occupation imbalance. We show that a method of moments readout of two quadratic spin observables J z 2 and J + 2 + J − 2 is globally optimal for Dicke state probes; i.e., the error saturates the QFI for all θ . In the lossy setting, we derive the time-inhomogeneous Markov process describing the effect of particle loss on TF states, showing that the method of moments readout of four at-most-quadratic spin observables is sufficient for globally optimal estimation of θ when two or more particles are lost. The analysis culminates in a numerical calculation of the QFI matrix for distributed MZ interferometry on the four-mode state | N 4 , N 4 , N 4 , N 4 〉 and its lossy counterparts, showing that an advantage for the estimation of any linear function of the local MZ phases θ 1 and θ 2 (compared to independent probing of the MZ phases by two copies of | N 4 , N 4 〉 ) appears when more than one particle is lost.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Minimum entropy filtering for a single output non-Gaussian stochastic system using state transformation

This paper presents a novel filter design for the single-output stochastic non-linear systems subjected to non-Gaussian noises and the proposed assumptions. Based on a state transformation, the unmeasurable states of the systems can be estimated where non-linear terms in the systems have been eliminated. It has been shown that the estimation error is linearly dynamical regarding to the presented vector-valued filter gain which can be optimised by minimising the entropy-based performance criterion. In addition, the convergence of the presented algorithm is analysed in mean-square sense and a numerical example is given to verify the effectiveness of the presented filtering algorithm. Meanwhile, the extended Kalman filter, unscented particle filter and minimum entropy filter are given for the comparisons of the filtering performance. Following the presented framework, some extensions of the presented filtering algorithm are discussed to indicate the flexibility of the filter design. The contribution of this paper can be summarised as establishing a novel minimum entropy filtering framework which consists of model transformation, entropy optimisation and convergence analysis.

42 ENGINEERING↗

Signal-preserving CMB component separation with machine learning

Analysis of microwave sky signals, such as the cosmic microwave background, often requires component separation using multifrequency methods, whereby different signals are isolated according to their different frequency behaviors. Many so-called blind methods, such as the internal linear combination (ILC), make minimal assumptions about the spatial distribution of the signal or contaminants, and only assume knowledge of the frequency dependence of the signal. The ILC produces a minimum-variance linear combination of the measured frequency maps. In the case of Gaussian, statistically isotropic fields, this is the optimal linear combination, as the variance is the only statistic of interest. However, in many cases the signal we wish to isolate, or the foregrounds we wish to remove, are non-Gaussian and/or statistically anisotropic (in particular for the case of Galactic foregrounds). In such cases, it is possible that machine learning (ML) techniques can be used to exploit the non-Gaussian features of the foregrounds and thereby improve component separation. However, many ML techniques require the use of complex, difficult-to-interpret operations on the data. We propose a hybrid method whereby we train an ML model using only combinations of the data that , and combine the resulting ML-predicted foreground estimate with the ILC solution to reduce the error from the ILC. We demonstrate our methods on simulations of extragalactic temperature and Galactic polarization foregrounds and show that our ML model can exploit non-Gaussian features, such as point sources and spatially varying spectral indices, to produce lower-variance maps than ILC—e.g., reducing the variance of the B-mode residual by factors of up to 5—while preserving the signal of interest in an unbiased manner. Moreover, we often find improved performance even when applying our ML technique to foreground models on which it was not trained. Published by the American Physical Society 2025

McCarthy, Fiona (ORCID:0000000253893565)↗