Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “probability and statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Self-calibrating the look-elsewhere effect: fast evaluation of the statistical significance using peak heights

In experiments where one searches a large parameter space for an anomaly, one often finds many spurious noise-induced peaks in the likelihood. This is known as the look-elsewhere effect, and must be corrected for when performing statistical analysis. Here, this paper introduces a method to calibrate the false alarm probability (FAP), or p-value, for a given dataset by considering the heights of the highest peaks in the likelihood. Specifically, we derive an equation relating the global p-value to the rank and height of local maxima. In the simplest form of self-calibration, the look-elsewhere-corrected $\chi^2$ of a physical peak is approximated by the $\chi^2$ of the peak minus the $\chi^2$ of the highest noise-induced peak, with accuracy improved by considering lower peaks. In contrast to alternative methods, this approach has negligible computational cost as peaks in the likelihood are a byproduct of every peak-search analysis. We apply to examples from astronomy, including planet detection, periodograms, and cosmology.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantifying Uncertainty in PV Energy Estimates Final Report

Uncertainty in PV energy estimates is "one of the most critical areas of lack of understanding" according to independent engineers, financiers, PV model developers, and other industry stakeholders. The primary problem is a lack of rigorous, transparent, widely accepted methods for quantifying uncertainty in energy production estimates. Uncertainty in energy production estimates arises from variability of the solar resource, inexact PV performance models and their parameters, and system reliability considerations. Uncertainty in annual energy production is frequently calculated for larger projects in order to quantify financial risk. Key statistics for energy, such as the P-values "P50" and "P90" (the annual energy values that are exceeded in future years with 50\% and 90\% probability, respectively) are used by financing institutions to calculate the repayment risk for the project. The current methods to estimate these statistics are typically proprietary, specialized, and involve significant post-processing of commercial performance model results. This black-box approach leads to inconsistent P-value estimates from different parties, which reduces investors' confidence in the results. Since the financial community bases its risk assessment on these estimates, reduced confidence increases perceived project risk, and consequently financing costs. The goal of this project was to establish a set of best practices for quantifying uncertainty in energy production estimates, including identifying what sources of uncertainty must be considered with clear definitions and metrics, determining which sources are the biggest drivers of uncertainty, and providing a computationally efficient framework for combining different sources of uncertainty that is flexible enough to accommodate substitutions of data or methods when better information is available. We engaged a wide set of stakeholders to ensure industry endorsement and adoption, and leveraged complementary projects investigating individual sources of uncertainty in great detail, as well as others' work that started down this path.

ENERGY PLANNING, POLICY, AND ECONOMY,SOLAR ENERGY↗

Single Gaussian process method for arbitrary tokamak regimes with a statistical analysis

Abstract Gaussian process regression is a Bayesian method for inferring profiles based on input data. The technique is increasing in popularity in the fusion community due to its many advantages over traditional fitting techniques including intrinsic uncertainty quantification and robustness to over-fitting. This work investigates the use of a new method, the change-point method, for handling the varying length scales found in different tokamak regimes. The use of the Student’s t-distribution for the Bayesian likelihood probability is also investigated and shown to be advantageous in providing good fits in profiles with many outliers. To compare different methods, synthetic data generated from analytic profiles is used to create a database enabling a quantitative statistical comparison of which methods perform the best. Using a full Bayesian approach with the change-point method, Matérn kernel for the prior probability, and Student’s t-distribution for the likelihood is shown to give the best results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Identifying Circumgalactic Medium Absorption in QSO Spectra: A Bayesian Approach

We present a study of candidate galaxy–absorber pairs for 43 low-redshift QSO sightlines (0.06 < z < 0.85) observed with the Hubble Space Telescope/Cosmic Origins Spectrograph that lie within the footprint of the Sloan Digital Sky Survey with a statistical approach to match absorbers with galaxies near the QSO lines of sight using only the SDSS Data Release 12 photometric data for the galaxies, including estimates of their redshifts. Our Bayesian methods combine the SDSS photometric information with measured properties of the circumgalactic medium to find the most probable galaxy match, if any, for each absorber in the line-of-sight QSO spectrum. We find ~630 candidate galaxy–absorber pairs using two different statistics. The methods are able to reproduce pairs reported in the targeted spectroscopic studies upon which we base the statistics at a rate of 72%. The properties of the galaxies comprising the candidate pairs have median redshift, luminosity, and stellar mass, all estimated from the photometric data, $z$ = 0.13, L = 0.1$L$ * , and $\mathrm{log}({M}_{* }/{M}_{\odot })=9.7$. The median impact parameter of the candidate pairs is ~430 kpc, or ~3.5 times the galaxy virial radius. The results are broadly consistent with the high Ly$α$ covering fraction out to this radius found in previous studies. In conclusion, this method of matching absorbers and galaxies can be used to prioritize targets for spectroscopic studies, and we present specific examples of promising systems for such follow-up.

79 ASTRONOMY AND ASTROPHYSICS↗

An update to the Sandia method for creating Typical Meteorological Years from a limited pool of calendar years

Typical Meteorological Years (TMYs) are essential for the efficient evaluation of energy system performance. Ideally, 30 years of weather data are required to generate TMYs, but significantly fewer years are typically available due to practical limitations. To address this issue, an update to the Sandia method was developed, referred to as the Argonne method, to create TMYs from a limited number of years. Furthermore, this method enhances candidate diversity by systematically shifting original candidate months forward or backward by specific days, creating an expanded pool of candidates. The effectiveness of the Argonne method was validated through statistical testing, comparison of monthly average weather parameters, and numerical simulations. The results demonstrate a high probability of identifying at least one shifted month whose cumulative distribution functions of weather parameters closely align with long-term distributions. In 67 % of all comparisons, the monthly average weather parameters in TMYs generated using the Argonne method exhibit better agreement with long-term averages than TMY3. Moreover, in 74 % of the 318 building simulation cases, the Argonne method outperforms TMY3 in estimating long-term average building heating and cooling demands. Therefore, the Argonne method effectively diversifies the candidate pool and produces typical years that provide more accurate estimations of long-term averages compared to TMY3 when only a limited pool of calendar years (10 years or fewer) is available.

Building energy modeling↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

Algorithm to extract direction in 2D discrete distributions and a continuous Frobenius norm

In this study, we present a novel algorithm for determining directionality in 2D distributions of discrete data. We compare a reference dataset with a known direction to a measured dataset with an unknown direction by the Frobenius norm of the difference (FND) to find the unknown direction. To generalize this concept, we develop a continuous Frobenius norm of the difference (CFND) as a continuous analog of the FND and derive its analytical expression. By relating fitted and normalized 2D Gaussian distributions, we show that the CFND approximates the FND, and we validate this relationship with computer simulations. We find that a first-order approximation of the CFND between two similar Gaussian distributions takes the form of an absolute sine function, offering a simple analytical form with potential for specialized applications in segmented inverse beta decay (IBD) neutrino detectors, astronomy, machine learning, and more. Although this method may easily extend to 3D scalar fields, our focus here is on 2D real-valued fields as it directly applies to directionality. Our methodology consists of modeling a 2D Gaussian distribution, binning the data into a histogram, and encoding it as a square matrix. Rotating this matrix around its geometric center and comparing it to a measured dataset using the FND gives us rotational data that we fit with an absolute sine function. The location of the minimum of this fit is the angle closest to the true angle of the direction in the measured dataset. We present the derivation and discuss initial applications of the CFND in our novel algorithm, demonstrating its success in approximating directionality in 2D distributions.

Data Analysis, Statistics and Probability (physics↗

A global significance evaluation method using simulated events

In High-Energy Physics experiments it is often necessary to evaluate the global statistical significance of apparent resonances observed in invariant mass spectra. One approach to determining significance is to use simulated events to find the probability of a random fluctuation in the background mimicking a real signal. As a high school summer project, we demonstrate a method with Monte Carlo simulated events to evaluate the global significance of a potential resonance with some assumptions. This method for determining significance is general and can be applied, with appropriate modification, to other resonances.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An implementation of neural simulation-based inference for parameter estimation in ATLAS

Neural simulation-based inference (NSBI) is a powerful class of machine-learning-based methods for statistical inference that naturally handles high-dimensional parameter estimation without the need to bin data into low-dimensional summary histograms. Such methods are promising for a range of measurements, including at the Large Hadron Collider, where no single observable may be optimal to scan over the entire theoretical phase space under consideration, or where binning data into histograms could result in a loss of sensitivity. This work develops a NSBI framework for statistical inference, using neural networks to estimate probability density ratios, which enables the application to a full-scale analysis. It incorporates a large number of systematic uncertainties, quantifies the uncertainty due to the finite number of events in training samples, develops a method to construct confidence intervals, and demonstrates a series of intermediate diagnostic checks that can be performed to validate the robustness of the method. As an example, the power and feasibility of the method are assessed on simulated data for a simplified version of an off-shell Higgs boson couplings measurement in the four-lepton final states. This approach represents an extension to the standard statistical methodology used by the experiments at the Large Hadron Collider, and can benefit many physics analyses.

frequentist statistics↗

Production of alternate realizations of DESI fiber assignment for unbiased clustering measurement in data and simulations

A critical requirement of spectroscopic large scale structure analyses is correcting for selection of which galaxies to observe from an isotropic target list. This selection is often limited by the hardware used to perform the survey which will impose angular constraints of simultaneously observable targets, requiring multiple passes to observe all of them. In SDSS this manifested solely as the collision of physical fibers and plugs placed in plates. In DESI, there is the additional constraint of the robotic positioner which controls each fiber being limited to a finite patrol radius. A number of approximate methods have previously been proposed to correct the galaxy clustering statistics for these effects, but these generally fail on small scales. To accurately correct the clustering we need to upweight pairs of galaxies based on the inverse probability that those pairs would be observed (Bianchi & Percival 2017). This paper details an implementation of that method to correct the Dark Energy Spectroscopic Instrument (DESI) survey for incompleteness. To calculate the required probabilities, we need a set of alternate realizations of DESI where we vary the relative priority of otherwise identical targets. These realizations take the form of alternate Merged Target Ledgers (AMTL), the files that link DESI observations and targets. We present the method used to generate these alternate realizations and how they are tracked forward in time using the real observational record and hardware status, propagating the survey as though the alternate orderings had been adopted. We detail the first applications of this method to the DESI One-Percent Survey (SV3) and the DESI year 1 data. We include evaluations of the pipeline outputs, estimation of survey completeness from this and other methods, and validation of the method using mock galaxy catalogs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Resonance Fitting for Nuclear Data Evaluation

Global and national efforts to deliver high-quality nuclear data to users have a wide-ranging impact, affecting applications in national security, reactor operations, basic science, medicine, and more. Cross section evaluation is a major part of this effort, combining theory and experimentation to produce recommended values and uncertainties for reaction probabilities. Resonance region evaluation is a specialized type of nuclear data evaluation that can require significant manual effort and months of time from expert scientists. In this article, non-convex non-linear optimization methods are combined with concepts of inferential statistics to infer a resonance model from experimental data in an automated manner that is not dependent on prior evaluation(s). This methodology aims to enhance the workflow of a resonance evaluator by minimizing time, effort, and the potential for bias from prior assumptions, while enhancing reproducibility and documentation, thereby addressing well-known challenges in the field.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Computing the Critical Temperature of the Affine-Transformed $D=3$ Ising Model Using Masked Autoregressive Flow

The simple Ising model provides a rich environment to build and study lattice field theories. As part of an ongoing project to construct a conformal field theory (CFT) on an arbitrarily curved manifold, in this work we develop methods to measure the critical temperature $β_c$ of the affine-transformed Ising model on the face-centered cubic (FCC) lattice. The main challenge in this endeavor is finding a computationally efficient and accurate method of interpolating and extrapolating Monte Carlo observables with respect to coupling coefficients and temperature. Herein, we compare two such methods. A traditional statistical approach uses the multiple histogram (MH) method, while a newer machine learning approach uses a masked autoregressive flow (MAF) to estimate the underlying probability density function of a set of observables. While the MH method is specifically designed to interpolate and extrapolate Monte Carlo observables, we find that MAF is a viable alternative for measuring $β_c$ with a computational cost that scales more favorably. Furthermore, we comment on additional advantages of MAF relevant to our work, such as extrapolating in system volume.

Svenson, Kai [Texas U.]↗

Generating high-resolution total canopy SIF emission from TROPOMI data: Algorithm and application

Solar-induced chlorophyll fluorescence (SIF) is a rapidly advancing front in modeling global terrestrial gross primary production (GPP). Canopy total SIF emissions (SIF total ) are mechanistically linked to the plant photosynthesis, and can be estimated from satellite observed SIF (SIF obs ) through radiative transfer modeling. However, the current satellite SIF obs and thus SIF total are available only at coarse spatial resolutions from several kilometers to tens of kilometers, inhibiting the application at fine spatial scales. Here, in this work, we proposed an algorithm to generate both global high-resolution SIF total (HSIF total ) and high-resolution SIF obs (HSIF obs ) at 1 km from low-resolution SIF obs (LSIF obs ) from the TROPOspheric Monitoring Instrument (TROPOMI), which has a spatial resolution at nadir of 3.5 km by 5.6–7 km. Our statistical method is based on the law of energy conservation and uses satellite derived fraction of absorbed photosynthetically active radiation, fluorescence efficiency, and the escape probability of fluorescence. We evaluated the accuracy of our HSIF total using the Orbiting Carbon Observatory-2 SIF (R 2 = 0.78). We found that the spatial resolution had clear effects on the relationship between HSIF total and GPP. We also compared HSIF total to 8-day averaged tower GPP from 135 flux sites and found that they were better correlated when HSIF total was averaged over a 1-km radius around the tower than when averaged over a larger radius. Our study provided a unique high-resolution HSIF total product, which will advance the estimation of GPP by extrapolating site-level relationships to the global scale.

54 ENVIRONMENTAL SCIENCES↗

Machine learning assisted bayesian inference of mix and hot-spot conditions in NIF implosions

Experiments on the National Ignition Facility (NIF) have provided clear evidence of ablator material mixing into the Hot-Spot, leading to degraded performance. However, inferring the amount of mix and Hot-Spot conditions from typical experimental observations (e.g. x-ray spectra and images) is highly challenging. Here, we have developed an analysis method that utilizes machine learning assisted Bayesian inference to find the probability distributions of the Hot-Spot and mix conditions. This approach uses a neural network, trained on an idealized 2-dimensional representation of the Hot-Spot and mix distribution, and Bayesian inference to find the statistical distributions of Hot-Spot conditions that provide a match with observations. We have tested this method with synthetic data from simulations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Numerical Analysis of High-Order Modes in SRF Resonators for Particle Accelerators

Over the past decades, superconducting technology has rapidly evolved towards high accelerating gradients and low surface resistance, making it possible to operate particle accelerators with high average beam currents and large duty factors. However, RF losses due to coherent excitation of the HOM become the limiting factor for these regimes. Unlike the cavity operating mode, which is tuned separately, the HOM parameters can significantly vary from one cavity to another due to finite mechanical tolerances during the manufacturing process. Thus, it is of utmost importance to know the HOM parameter spread in advance in order to predict unexpected cryogenic losses, overheating of beam line components and maintain stable beam dynamics. In this paper, we present a method for generating cavity geometry with an arbitrary spread of mechanical imperfections and numerically evaluating HOM statistics. Knowing the spread of HOM parameters, we calculated the probability of resonant HOM losses in SRF accelerating cavities used in CW beam current machines such as the PIP-II and LCLS-II linacs, as well as for the SRF crab-cavity for the ILC project. Finally, we present experimental results of HOM spectra measurements in hundreds of 1.3 GHz cavities installed in LCLS-II cryomodules. Studying the effects of HOM excitation results in specifications of the SRF cavity and cryomodule and can significantly impact the efficiency and reliability of the machine operation.

Lunin, Andrei [Fermilab] (ORCID:0000000290096792)↗

Statistical analysis for the neutrinoless double-β-decay matrix element of 48 Ca

Neutrinoless double-β-decay (0⁢vββ) nuclear matrix elements (NME) are the object of many theoretical calculation methods, and are very important for analysis and guidance of a large number of experimental efforts. However, there are large discrepancies between the NME values provided by different methods. Here, in this paper, we propose a statistical analysis of the 48 Ca 0v⁢ββ NME using the interacting shell model, emphasizing the range of the NME probable values and their correlations with observables that can be obtained from the existing nuclear data. Based on this statistical analysis with three independent effective Hamiltonians, we propose a common probability distribution function for the 0⁢vββ NME, which has a range of (0.45–0.95) at 90% confidence level, and a mean value of 0.68.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗