Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “applied statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Quantifying uncertainties and correlations in the nuclear-matter equation of state

We perform statistically rigorous uncertainty quantification (UQ) for chiral effective field theory (χ EFT) applied to infinite nuclear matter up to twice nuclear saturation density. The equation of state (EOS) is based on high-order many-body perturbation theory calculations with nucleon-nucleon and three-nucleon interactions up to fourth order in the χ EFT expansion. From these calculations our newly developed Bayesian machine-learning approach extracts the size and smoothness properties of the correlated EFT truncation error. Furthermore, we then propose a novel extension that uses multitask machine learning to reveal correlations between the EOS at different proton fractions. The inferred in-medium χ EFT breakdown scale in pure neutron matter and symmetric nuclear matter is consistent with that from free-space nucleon-nucleon scattering. These significant advances allow us to provide posterior distributions for the nuclear saturation point and propagate theoretical uncertainties to derived quantities: the pressure and incompressibility of symmetric nuclear matter, the nuclear symmetry energy, and its derivative. Our results, which are validated by statistical diagnostics, demonstrate that an understanding of truncation-error correlations between different densities and different observables is crucial for reliable UQ. The methods developed here are publicly available as annotated Jupyter notebooks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Dark Energy Survey Year 3 Results: Cosmological constraints from second- and third-order shear statistics

Here, we present a cosmological analysis of the third-order aperture mass statistic using Dark Energy Survey Year 3 (DES Y3) data. We perform a complete tomographic measurement of the three-point correlation function of the Y3 weak lensing shape catalog with the four fiducial source redshift bins. Building upon our companion methodology paper, we apply a pipeline that combines the two-point function ξ ± with the mass aperture skewness statistic ⟨ M ap 3 ⟩ , which is an efficient compression of the full shear three-point function. We use a suite of simulated shear maps to obtain a joint covariance matrix. By jointly analyzing ξ ± and ⟨ M ap 3 ⟩ measured from DES Y3 data with a Λ CDM model, we find S 8 = 0.780 ± 0.015 and Ω m = 0.26 6 - 0.040 + 0.039 , yielding 111% of figure-of-merit improvement in the Ω m - S 8 plane relative to ξ ± alone, consistent with expectations from simulated likelihood analyses. With a w CDM model, we find S 8 = 0.74 9 - 0.026 + 0.027 and w 0 = - 1.39 ± 0.31 , which gives an improvement of 22% on the joint S 8 - w 0 constraint. Our results are consistent with w 0 = - 1 . Our new constraints are compared to CMB data from the Planck satellite, and we find that with the inclusion of ⟨ M ap 3 ⟩ the existing tension between the datasets is at the level of 2.3 σ . We show that the third-order statistic enables us to self-calibrate the mean photometric redshift uncertainty parameter of the highest redshift bin with little degradation in the figure of merit. Our results demonstrate the constraining power of higher-order lensing statistics and establish ⟨ M ap 3 ⟩ as a practical observable for joint analyses in current and future surveys.

Gomes, R. C. H. [University of Pennsylvania] (ORCI↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Bayesian Adaptive Polynomial Chaos Expansions

Polynomial chaos expansions (PCEs) are widely used for uncertainty quantification (UQ) tasks, particularly in the applied mathematics community. However, PCE has received comparatively less attention in the statistics literature, and fully Bayesian formulations remain rare—especially with implementations in R. Motivated by the success of adaptive Bayesian machine learning models such as BART, BASS and BPPR, we develop a new fully Bayesian adaptive PCE method with an efficient and accessible R implementation: khaos. Our approach includes a novel proposal distribution that enables data-driven interaction selection and supports a modified g-prior tailored to PCE structure. Through simulation studies and real-world UQ applications, we demonstrate that the Bayesian adaptive PCE provides competitive performance for surrogate modeling, global sensitivity analysis and ordinal regression tasks.

97 MATHEMATICS AND COMPUTING↗

Multi-fidelity Uncertainty Quantification for Homogenization Problems in Structure-Property Relationships from Crystal Plasticity Finite Elements

Crystal plasticity finite element method (CPFEM) has been an integrated computational materials engineering (ICME) workhorse to study materials behaviors and structure-property relationships for the last few decades. These relations are mappings from the microstructure space to the materials properties space. Due to the stochastic and random nature of microstructures, there is always some uncertainty associated with materials properties, for example, in homogenized stress-strain curves. For critical applications with strong reliability needs, it is often desirable to quantify the microstructure-induced uncertainty in the context of structure-property relationships. However, this uncertainty quantification (UQ) problem often incurs a large computational cost because many statistically equivalent representative volume elements (SERVEs) are needed. In this article, we apply a multi-level Monte Carlo (MLMC) method to CPFEM to study the uncertainty in stress-strain curves, given an ensemble of SERVEs at multiple mesh resolutions. By using the information at coarse meshes, we show that it is possible to approximate the response at fine meshes with a much reduced computational cost. We focus on problems where the model output is multi-dimensional, which requires us to track multiple quantities of interest (QoIs) at the same time. In conclusion, our numerical results show that MLMC can accelerate UQ tasks around 2.23x, compared to the classical Monte Carlo (MC) method, which is widely known as ensemble average in the CPFEM literature.

36 MATERIALS SCIENCE↗

Stochastic Unit Commitment: Model Reduction via Learning

As weather-dependent renewable generation increases its share in the generation mix of most electric energy systems, a stochastic unit commitment becomes the natural day-ahead scheduling tool. However, such a tool is generally computationally intractable if a detailed uncertainty description is considered. Taking this into account, we proposed a learning method to make the stochastic unit commitment problem tractable. Here, recent advances in statistical learning and machine learning to address optimization problems can be advantageously applied to the rather intractable stochastic unit commitment problem. Considering these advances, we explore simple learning techniques to drastically reduce the size of a stochastic unit commitment problem without significantly altering its optimal solution. The considered stochastic unit commitment problem is formulated as a two-stage stochastic programming problem. The first stage represents commitment decisions, while the second one represents the operation conditions under different scenarios. Taking into account historical solved instances (or proxies for them), we reduce the size (measured by numbers of constraints and variables) of the stochastic unit commitment problem by (i) fixing unchanged binary variables and by (ii) eliminating inactive inequality constraints. Our numerical results show that the reduced problem generally requires significantly less time to solve while obtaining high-quality solutions, which are very close to or indistinguishable from the one obtained by solving the original problem. We use an Illinois 200-bus system to illustrate and characterize the performance of the proposed problem-reduction method.

42 ENGINEERING↗

Probing the role of solids loading and mix procedure on the properties of acoustically mixed materials for additive manufacturing

We report resonant acoustic mixing has been of particular interest for use in additive manufacturing since viscous, solids-loaded materials can be difficult to mix and inhomogeneity has adverse effects on print quality. In this study, we detail a method to iterate through different formulations and mix procedures and assess mixture quality. The approach utilizes a constant pressure-driven flow test to collect statistics on flow rate through a standard geometry. This test is first applied to well mixed formulations containing particulate solids and find that the results are sensitive to viscosity produced by solids content. We then consider the formulations at various stages of mixing and find that the spread in the volume measurements is indicative of the mixing quality; poorly mixed material yields measurements with wide distributions. We believe this testing approach can be useful when screening new formulations, developing mixing processes, or for quality control.

36 MATERIALS SCIENCE↗

Connectivity-informed drainage network generation using deep convolution generative adversarial networks

Abstract Stochastic network modeling is often limited by high computational costs to generate a large number of networks enough for meaningful statistical evaluation. In this study, Deep Convolutional Generative Adversarial Networks (DCGANs) were applied to quickly reproduce drainage networks from the already generated network samples without repetitive long modeling of the stochastic network model, Gibb’s model. In particular, we developed a novel connectivity-informed method that converts the drainage network images to the directional information of flow on each node of the drainage network, and then transforms it into multiple binary layers where the connectivity constraints between nodes in the drainage network are stored. DCGANs trained with three different types of training samples were compared; (1) original drainage network images, (2) their corresponding directional information only, and (3) the connectivity-informed directional information. A comparison of generated images demonstrated that the novel connectivity-informed method outperformed the other two methods by training DCGANs more effectively and better reproducing accurate drainage networks due to its compact representation of the network complexity and connectivity. This work highlights that DCGANs can be applicable for high contrast images common in earth and material sciences where the network, fractures, and other high contrast features are important.

42 ENGINEERING↗

Using active matter to introduce spatial heterogeneity to the susceptible infected recovered model of epidemic spreading

Abstract The widely used susceptible-infected-recovered (S-I-R) epidemic model assumes a uniform, well-mixed population, and incorporation of spatial heterogeneities remains a major challenge. Understanding failures of the mixing assumption is important for designing effective disease mitigation approaches. We combine a run-and-tumble self-propelled active matter system with an S-I-R model to capture the effects of spatial disorder. Working in the motility-induced phase separation regime both with and without quenched disorder, we find two epidemic regimes. For low transmissibility, quenched disorder lowers the frequency of epidemics and increases their average duration. For high transmissibility, the epidemic spreads as a front and the epidemic curves are less sensitive to quenched disorder; however, within this regime it is possible for quenched disorder to enhance the contagion by creating regions of higher particle densities. We discuss how this system could be realized using artificial swimmers with mobile optical traps operated on a feedback loop.

60 APPLIED LIFE SCIENCES↗

A hybrid neural architecture: Online attosecond x-ray characterization

The emergence of high-repetition-rate x-ray free-electron lasers (XFELs), such as SLAC’s LCLS-II, serves as our canonical example for autonomous controls that necessitate high-throughput diagnostics paired with streaming computational pipelines capable of single-shot analysis with extremely low latency. We present the deterministic characterization with an integrated parallelizable hybrid resolver architecture, a hybrid machine learning framework designed for fast, accurate analysis of XFEL diagnostics using angular streaking-based sinogram images. This architecture integrates convolutional neural networks and bidirectional long short-term memory models to denoise input, identify x-ray sub-spike features, and extract sub-spike relative delays with sub-30 attosecond temporal resolution. Deployed on low-latency hardware, it achieves over 10 kHz throughput with 168.3 μs inference latency, indicating scalability to 14 kHz with field-programmable gate array integration. By transforming regression tasks into classification problems and leveraging optimized error encoding, we achieve high precision with low-latency performance that is critical for real-time streaming event selection and experimental control feedback signals. This represents a key development in real-time control pipelines for next-generation autonomous science, generally, and high repetition-rate x-ray experiments in particular.

Accelerator Physics (physics.acc-ph)↗

The Atacama Cosmology Telescope: Large-scale velocity reconstruction with the kinematic Sunyaev-Zel'dovich effect and DESI LRGs

The kinematic Sunyaev-Zel'dovich (kSZ) effect induces a non-zero density-density-temperature bispectrum, which we can use to reconstruct the large-scale velocity field from a combination of cosmic microwave background (CMB) and galaxy density measurements, in a procedure known as “kSZ velocity reconstruction”. This method has been forecast to constrain large-scale modes with future galaxy and CMB surveys, improving their measurement beyond what is possible with the galaxy surveys alone. Such measurements will enable tighter constraints on large-scale signals such as primordial non-Gaussianity, deviations from homogeneity, and modified gravity. In this work, we demonstrate a statistically significant measurement of kSZ velocity reconstruction for the first time, by applying quadratic estimators to the combination of the ACT DR6 CMB+kSZ map and the DESI LRG galaxies (with photometric redshifts) in order to reconstruct the velocity field. We do so using a formalism appropriate for the 2-dimensional projected galaxy fields that we use, which naturally incorporates the curved-sky effects important on the largest scales. We find evidence for the signal by cross-correlating with an external estimate of the velocity field from the spectroscopic BOSS survey and rejecting the null (no-kSZ) hypothesis at 3.8σ. Our work presents a first step towards the use of this observable for cosmological analyses.

Sunyaev-Zeldovich effect↗

QUIJOTE scientific results – IV. A northern sky survey in intensity and polarization at 10–20 GHz with the multifrequency instrument

ABSTRACT We present QUIJOTE intensity and polarization maps in four frequency bands centred around 11, 13, 17, and 19 GHz, and covering approximately 29 000 deg2, including most of the northern sky region. These maps result from 9000 h of observations taken between May 2013 and June 2018 with the first QUIJOTE multifrequency instrument (MFI), and have angular resolutions of around 1°, and sensitivities in polarization within the range 35–40 µK per 1° beam, being a factor ∼2–4 worse in intensity. We discuss the data processing pipeline employed, and the basic characteristics of the maps in terms of real space statistics and angular power spectra. A number of validation tests have been applied to characterize the accuracy of the calibration and the residual level of systematic effects, finding a conservative overall calibration uncertainty of 5 per cent. We also discuss flux densities for four bright celestial sources (Tau A, Cas A, Cyg A, and 3C274), which are often used as calibrators at microwave frequencies. The polarization signal in our maps is dominated by synchrotron emission. The distribution of spectral index values between the 11 GHz and WMAP 23 GHz map peaks at β = −3.09 with a standard deviation of 0.14. The measured BB/EE ratio at scales of ℓ = 80 is 0.26 ± 0.07 for a Galactic cut |b| > 10°. We find a positive TE correlation for 11 GHz at large angular scales (ℓ ≲ 50), while the EB and TB signals are consistent with zero in the multipole range 30 ≲ ℓ ≲ 150. The maps discussed in this paper are publicly available.

Astronomy & Astrophysics↗

First test of the consistency relation for the large-scale structure using the anisotropic three-point correlation function of BOSS DR12 galaxies

ABSTRACT We present, for the first time, an observational test of the consistency relation for the large-scale structure (LSS) of the Universe through a joint analysis of the anisotropic two- and three-point correlation functions (2PCF and 3PCF) of galaxies. We parameterize the breakdown of the LSS consistency relation in the squeezed limit by Es, which represents the ratio of the coefficients of the shift terms in the second-order density and velocity fluctuations. Es ≠ 1 is a sufficient condition under which the LSS consistency relation is violated. A novel aspect of this work is that we constrain Es by obtaining information about the non-linear velocity field from the quadrupole component of the 3PCF without taking the squeezed limit. Using the galaxy catalogues in the Baryon Oscillation Spectroscopic Survey (BOSS) Data Release 12, we obtain $E_{\rm s} = -0.92_{-3.26}^{+3.13}$, indicating that there is no violation of the LSS consistency relation in our analysis within the statistical errors. Our parameterization is general enough that our constraint can be applied to a wide range of theories, such as multicomponent fluids, modified gravity theories, and their associated galaxy bias effects. Our analysis opens a new observational window to test the fundamental physics using the anisotropic higher-order correlation functions of galaxy clustering.

79 ASTRONOMY AND ASTROPHYSICS↗

Characterization of single-shot attosecond pulses with angular streaking photoelectron spectra

Most of the traditional attosecond pulse retrieval algorithms are based on a so-called attosecond streak camera technique, in which the momentum of the electron is shifted by an amount depending on the relative time delay between the attosecond pulse and the streaking infrared pulse. Thus, temporal information of the attosecond pulse is encoded in the amount of momentum shift in the streaked photoelectron momentum spectrogram S(p,τ), where p is the momentum of the electron along the polarization direction and τ is the time delay. An iterative algorithm is then employed to reconstruct the attosecond pulse from the streaking spectrogram. This method, however, cannot be applied to attosecond pulses generated from free-electron x-ray lasers where each single shot is different and stochastic in time. However, using a circularly polarized infrared laser as the streaking field, a two (or three)-dimensional angular streaking electron spectrum can be used to retrieve attosecond pulses for each shot, as well as the time delay with respect to the circularly polarized IR field. Here we show that a retrieval algorithm previously developed for the traditional streaking spectrogram can be modified to efficiently characterize single-shot attosecond pulses. The methods have been applied to retrieve 188 single shots from recent experiments. We analyze the statistical behavior of these 188 pulses in terms of pulse duration, bandwidth, pulse peak energy, and time delay with respect to the IR field. Furthermore, the retrieval algorithm is efficient and can be easily used to characterize a large number of shots in future experiments for attosecond pulses at free-electron x-ray laser facilities.

74 ATOMIC AND MOLECULAR PHYSICS↗

Accurate estimation of angular power spectra for maps with correlated masks

A common procedure when analyzing maps of the cosmic microwave background (CMB) or other cosmological signals is the need to remove ("mask") regions of the maps that are heavily contaminated, e.g., by non-cosmological foreground emission. After applying such a mask, one must account for its effect when inferring statistical properties of interest, such as the angular power spectrum of the field in the original map. A widely used approach to correct for such mask-induced effects was presented by Hivon et al. (2002), now widely known as the "MASTER" formalism. However, it is often the case that the map and mask are correlated in some way, such as point source masks used in CMB analyses, which have nonzero correlation with CMB secondary anisotropy fields and other mm-wave sky signals. In such situations, the MASTER approach gives biased results, as it assumes that the unmasked map and mask have zero correlation. While such effects have been discussed before with regard to specific physical models, here we derive a completely general formalism for any case where the map and mask are correlated. We show that our result ("reMASTERed") reconstructs ensemble-averaged angular power spectra to effectively exact precision, with significant improvements over traditional estimators for cases where the map and mask are correlated. An important consequence of our result is that for maps with correlated masks, it is no longer possible to invert a simple equation to obtain the true power spectrum from the observed (masked) power spectrum. Instead, our result necessitates the use of forward modeling from theory space into the observable domain of the masked power spectrum. We publicly release our software implementation of these results.

79 ASTRONOMY AND ASTROPHYSICS↗

Λ Baryon Production in ν¯µ Interactions in the MicroBooNE Detector

The Cabibbo suppressed production of $\Lambda$ baryons in anti-neutrino interactions with nuclei is a rare process that is yet to be measured with a modern neutrino detector with automated reconstruction. The cross section for this process is sensitive to a number of unique nuclear effects, most notably the secondary interactions of the produced hyperon while attempting to escape from the nucleus. Other interactions within the nuclear remnant can impact the estimation of neutrino energy in oscillation measurements, and thus an accurate description of the nuclear environment is required. The strangeness violating hyperon production process is only available to anti-neutrinos. The model of this interaction is implemented into the NuWro neutrino interaction Monte Carlo simulation, and some predictions are presented, focusing on the role of nuclear effects. This model introduces a hyperon-nucleus potential, which calculations from hypernuclear theory permit to be strongly repulsive in th e case of $\Sigma$ baryons. The presence of this potential is found to sculpt the shape of the differential cross section in some variables. The MicroBooNE detector will be described, followed by a description of a measurement of the flux averaged, restricted phase space cross section of Cabibbo suppressed $\Lambda$ baryon production. A sophisticated event selection is employed, as a very large quantity of background neutrino interactions must be removed to perform the measurement with any sensitivity. This selection introduces some novel techniques such as the island finding method, and achieves a background reduction of $\sim 10^6$, with an efficiency of around 7\%. The calculation of the systematic uncertainties will be explained, including two procedures explored to handle sources of background with extremely poor simulation statistics: an in-situ constraint using data from sidebands, and a visual inspection of the data and simulation to remove the troublesome background events. The sensitivity to the $\Lambda$ baryon production cross section is calculated in the form of Bayesian posterior probability distributions, combining the systematic uncertainties with data and simulation statistical uncertainties. As a rare process, the statistical uncertainties are highly non-Gaussian, and the Bayesian approach is applied to include the full shapes of these uncertainties. Data corresponding to $2.2 \times 10^{20}$ protons on target of neutrino mode running and $4.9 \times 10^{20}$ protons on target of anti-neutrino running is analysed. When the data was unblinded, five $\Lambda$ production candidates were selected from the data, consistent with the MC simulation prediction of $5.3 \pm 1.1$ events. The final estimated cross section is $1.8^{+2.0}_{-1.6} \times 10^{-40}$cm$^2/$Ar when employing the sideband constraint procedure. A similar result of $2.0^{+2.2}_{-1.8} \times 10^{-40}$cm$^2/$Ar is obtained when performing the visual scan instead. The methods used in t his analysis are intended to be easily exported to other LArTPC detectors such as the Short Baseline Near Detector.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Smokescreen: A Python package for data vector blinding and encryption in cosmological analyses

Smokescreen is an open-source Python library for data-vector concealment (blinding) in cosmological analyses. Data-vector blinding works by applying cosmology-dependent shifts to the observed data vector, moving it away from the true cosmological signal without affecting its statistical properties, so that analysts cannot infer the true result until the analysis is frozen and the blinding is lifted. The package computes these shifts using Firecrown likelihoods applied to data vectors stored in the SACC format, ensuring that the theoretical model used for blinding is identical to that used for inference whilst remaining agnostic to the specific observable being blinded. To prevent accidental unblinding, the original SACC file, containing the true cosmology, is encrypted. Although developed for the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), Smokescreen is applicable to any experiment using Firecrown likelihoods and the SACC data format.

Loureiro, Arthur [Stockholm U., OKC; Imperial Coll↗

Towards a Comparative Assessment of Data-Driven Process Models in Health Information Technology

Process mining for conformance analysis focuses on comparing a reference process model against a data-driven process model that is generated via log files from information technology systems. While this approach is helpful when there is an existing process model in an organization, it leaves the question of what to do in the absence of a complete reference process model unanswered. In this paper, we present a comparative assessment approach that combines process mining, process mapping for dimensionality reduction, and statistical analysis. Our goal is to find similarities and dissimilarities in data-driven process models among U.S. Veterans Health Administration (VHA) facilities to assess process conformance among different healthcare facilities, which can help assess the standardization of care. We illustrate our approach by applying it to two clinical radiology order process models generated by two similar facilities. Our results demonstrate statistical similarities in the standardization of care among those two facilities.

Klasky, Hilda↗