Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical sampling techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

ADDGALS: Simulated Sky Catalogs for Wide Field Galaxy Surveys

Abstract We present a method for creating simulated galaxy catalogs with realistic galaxy luminosities, broadband colors, and projected clustering over large cosmic volumes. The technique, denoted Addgals (Adding Density Dependent GAlaxies to Lightcone Simulations), uses an empirical approach to place galaxies within lightcone outputs of cosmological simulations. It can be applied to significantly lower-resolution simulations than those required for commonly used methods such as halo occupation distributions, subhalo abundance matching, and semi-analytic models, while still accurately reproducing projected galaxy clustering statistics down to scales of r ∼ 100 h −1 kpc . We show that Addgals catalogs reproduce several statistical properties of the galaxy distribution as measured by the Sloan Digital Sky Survey (SDSS) main galaxy sample, including galaxy number densities, observed magnitude and color distributions, as well as luminosity- and color-dependent clustering. We also compare to cluster–galaxy cross correlations, where we find significant discrepancies with measurements from SDSS that are likely linked to artificial subhalo disruption in the simulations. Applications of this model to simulations of deep wide-area photometric surveys, including modeling weak-lensing statistics, photometric redshifts, and galaxy cluster finding, are presented in DeRose et al., and an application to a full cosmology analysis of Dark Energy Survey (DES) Year 3 like data is presented in DeRose et al. We plan to publicly release a 10,313 square degree catalog constructed using Addgals with magnitudes appropriate for several existing and planned surveys, including SDSS, DES, VISTA, Wide-field Infrared Survey Explorer, and Rubin Observatory’s Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS↗

A review of imputation strategies for isobaric labeling-based shotgun proteomics

The throughput efficiency and increased depth of coverage provided by isobaric-labeled proteomics measurements have led to increased usage of these techniques. However, the structure of missing data is uniquely different than unlabeled studies. In this review, we compare the efficacy of nine imputation methods on a CPTAC proteomics iTRAQ dataset. Imputation methods were evaluated with regard to accuracy, variability, statistical hypothesis test inference and run time over datasets consisting of varying number of iTRAQ plexes and percentages of missing data. In general, expectation maximization and random forest imputation methods yielded the best performances, and constant-based methods performed poorly consistently across all dataset sizes and percentages of missing values. For datasets with small sample sizes and higher percentages of missing data, results indicate that statistical inference with no imputation may be preferable. Based on the findings in this review, there are core imputation methods that perform higher for isobaric-labeled proteomics data, but great care and consideration as to whether imputation should be used should be given for datasets comprised of a small number of samples, as well as to factors such as computational time and reproducibility of imputation values.

Bramer, Lisa M.↗

Red Noise–based False Alarm Thresholds for Astrophysical Periodograms via Whittle’s Approximation to the Likelihood

Astronomers who search for periodic signals using Lomb–Scargle periodograms rely on false alarm level (FAL) estimates to identify statistically significant peaks. Although FALs are often calculated from white noise models, many astronomical time series suffer from red noise. Prewhitening is a statistical technique in which a continuum model is subtracted from the log power spectrum estimate, after which the observer can proceed with a white-noise treatment. Here we present a prewhitening-based method of calculating frequency-dependent FALs. We fit power laws and autoregressive models of order 1 to each Lomb–Scargle periodogram by minimizing the Whittle approximation to the negative log-likelihood (NLL), then calculate FALs based on the best-fit model power spectrum. Our technique is a novel extension of the Whittle NLL to datasets with uneven time sampling. We demonstrate FAL calculations using observations of α Cen B, GJ 581, HD 192310, synthetic data from the radial velocity (RV) fitting challenge, and Kepler observations of a differential rotator. The Kepler data analysis shows that only true rotation signals are detected by red noise FALs, while white noise FALs suggest all spurious peaks in the low-frequency range are significant. A high-frequency sinusoid injected into α Cen B logR$'$ HK observations exceeds the 1% red noise FAL despite having only 8.9% of the power of the dominant rotation signal. In a periodogram of HD 192310 RVs, peaks associated with differential rotation and planets are detected against the 5% red noise FAL without iterative model fitting or subtraction. The software for calculating red noise–based FALs is available on GitHub.

Astrostatistics (1882)↗

AutoEnRichness: A hybrid empirical and analytical approach for estimating the richness of galaxy clusters

ABSTRACT We introduce AutoEnRichness, a hybrid approach that combines empirical and analytical strategies to determine the richness of galaxy clusters (in the redshift range of 0.1 ≤ z ≤ 0.35) using photometry data from the Sloan Digital Sky Survey Data Release 16, where cluster richness can be used as a proxy for cluster mass. In order to reliably estimate cluster richness, it is vital that the background subtraction is as accurate as possible when distinguishing cluster and field galaxies to mitigate severe contamination. AutoEnRichness is comprised of a multistage machine learning algorithm that performs background subtraction of interloping field galaxies along the cluster line of sight and a conventional luminosity distribution fitting approach that estimates cluster richness based only on the number of galaxies within a magnitude range and search area. In this proof-of-concept study, we obtain a balanced accuracy of 83.20 per cent when distinguishing between cluster and field galaxies as well as a median absolute percentage error of 33.50 per cent between our estimated cluster richnesses and known cluster richnesses within r200. In the future, we aim for AutoEnRichness to be applied on upcoming large-scale optical surveys, such as the Legacy Survey of Space and Time and Euclid, to estimate the richness of a large sample of galaxy groups and clusters from across the halo mass function. This would advance our overall understanding of galaxy evolution within overdense environments as well as enable cosmological parameters to be further constrained.

79 ASTRONOMY AND ASTROPHYSICS↗

Error statistics and scalability of quantum error mitigation formulas

Quantum computing promises advantages over classical computing in many problems. Nevertheless, noise in quantum devices prevents most quantum algorithms from achieving the quantum advantage. Quantum error mitigation provides a variety of protocols to handle such noise using minimal qubit resources. While some of those protocols have been implemented in experiments for a few qubits, it remains unclear whether error mitigation will be effective in quantum circuits with tens to hundreds of qubits. In this paper, we apply statistics principles to quantum error mitigation and analyse the scaling behaviour of its intrinsic error. We find that the error increases linearly O(ϵN) with the gate number N before mitigation and sublinearly O(ϵ'N γ ) after mitigation, where γ ≈ 0.5, ϵ is the error rate of a quantum gate, and ϵ' is a protocol-dependent factor. The $\sqrt{N}$ scaling is a consequence of the law of large numbers, and it indicates that error mitigation can suppress the error by a larger factor in larger circuits. We propose the importance Clifford sampling as a key technique for error mitigation in large circuits to obtain this result.

97 MATHEMATICS AND COMPUTING↗

One Galaxy Sample to Rule Them All: Halo Occupation Distribution Modeling of DES Year 3 Source Galaxies

Abstract For the joint analysis of second-order weak-lensing and galaxy clustering statistics, so-called 3 × 2 analyses, the selection and characterization of optimal galaxy samples is a major area of research. One promising choice is to use the same galaxy sample as lenses and sources, which reduces the systematics parameter space that describes the uncertainties related to galaxy samples. Such a “lens-equal-source” analysis significantly improves the self-calibration of photo- z systematics, leading to improved cosmological constraints. With the aim of enabling a lens-equal-source analysis on small scales, we investigate the halo–galaxy connection of DES Year 3 source galaxies. We develop a technique to construct mock source galaxy populations by matching COSMOS/UltraVISTA photometry to U niverse M achine galaxies. These mocks predict a source halo occupation distribution (HOD) that exhibits significant redshift evolution, nontrivial central incompleteness, and galaxy assembly bias. We produce multiple realizations of mock source galaxies drawn from the U niverse M achine posterior, with added uncertainties in the measured Dark Energy Survey photometry and galaxy shapes. We fit a modified HOD formalism to these realizations to produce priors on the galaxy–halo connection for cosmological analyses. We additionally train an emulator that predicts this HOD to ∼2% accuracy from redshift z = 0.1−1.3 that models the dependence of this HOD on (1) observational uncertainties in galaxy size and photometry and (2) uncertainties in the U niverse M achine predictions.

Salcedo, Andrés N. (ORCID:000000031420527X)↗

Spatiotemporal Learning in Power Modules: Wavelet-Enhanced Forecasting of Thermomechanical Degradation

Detecting internal defects in power electronics packages is critical for their performance and reliability, especially under extreme operating conditions, as these defects can lead to catastrophic failure if not properly addressed. Confocal scanning acoustic microscopy (C-SAM) plays a key role in the nondestructive evaluation of bond layer degradation within a power electronics package by detecting defects such as delamination, voids, and cracks. However, accurately quantifying and predicting these defects from C-SAM images remains a significant challenge due to the low noise-to-signal ratio, which typically arises from both imaging process and bond patterns itself. In this paper, we explore machine learning strategies for processing C-SAM images and providing predictive models of defect growth. We use C-SAM images of sintered copper and sintered silver samples, which are obtained under accelerated thermal experiments, as the representative dataset for our study. We investigate the effect of Fourier transforms and wavelet transforms on these datasets to remove high-frequency noise and address noise across multiple scales with histogram equalization to enhance the contrast and improve the visibility of defects. As a result, defect boundaries can be clearly distinguished, enabling more accurate tracking of their growth over time. We then employ different time-series forecasting algorithms on the denoised images to formulate an image-based lifetime prediction model. Statistical models and deep-learning techniques are trained on images obtained in the early stages of thermal shock, and defect growth in the later stages is predicted. Our work serves as a preliminary attempt to improve the accuracy of lifetime prediction models of power electronics packages, which is critical under extreme operating environments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

TRISO SiC Failure Probability for Reactivity Initiated Accidents in High-Temperature Gas-Cooled Reactors

This work analyzes the failure process of the silicon carbide (SiC) layer in tristructural isotropic (TRISO) during reactivity-initiated accident scenarios for a high-temperature gas-cooled reactor (HTGR) with BISON. Two cases are considered—a group control rod withdrawal (CRW) and a control rod ejection (CRE)—reproduced from a previous study. Failure probability is modeled using Weibull statistics, and worst-case scenario Weibull parameters are adopted to simulate the envelopes in BISON with a one-dimensional TRISO model. CRW scenario results are characterized by higher values of maximum energy deposition and final temperature and volumetric strain with respect to the CRE ones, but the latter have remarkably higher SiC failure probability, mainly due to the offset in strain rates between the two cases. This work also confirms the validity and conservatism of the performance envelopes produced in a previous work by replicating the envelope formulation using RELAP5-3D and RAVEN with a different sampling technique and obtaining consistent results. A sensitivity analysis using the Sobol variance decomposition method on SiC failure probability is then performed involving a set of inputs on both CRW and CRE. The two most important parameters are Weibull modulus and characteristic stress, and their relative importance depends on the specific case. The proposed interpretation of the results is that both energy deposition and strain rate influence the relative degree of importance of the failure parameters. Computation of 95% confidence intervals around worst-case scenario SiC failure probability values is also carried out for four different sets of Weibull parameters. Heren a new criterion for SiC TRISO quality classification built upon safety-based ranges of Weibull parameters is proposed to be integrated in future Fuel-Production Quality Assurance Plans.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Breeding Realistic D‐Brane Models

Abstract Intersecting branes provide a useful mechanism to construct particle physics models from string theory with a wide variety of desirable characteristics. The landscape of such models can be enormous, and navigating towards regions which are most phenomenologically interesting is potentially challenging. Machine learning techniques can be used to efficiently construct large numbers of consistent and phenomenologically desirable models. In this work we phrase the problem of finding consistent intersecting D‐brane models in terms of genetic algorithms, which mimic natural selection to evolve a population collectively towards optimal solutions. For a four‐dimensional supersymmetric type IIA orientifold with intersecting D6‐branes, we demonstrate that unique, fully consistent models can be easily constructed, and, by a judicious choice of search environment and hyper‐parameters, of the found models contain the desired Standard Model gauge group factor. Having a sizable sample allows us to draw some preliminary landscape statistics of intersecting brane models both with and without the restriction of having the Standard Model gauge factor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Probing the PeV region in the astrophysical neutrino spectrum using 𝜈 𝜇 from the Southern sky

IceCube has observed a diffuse astrophysical neutrino flux over the energy region from a few TeV to a few PeV. At PeV energies, the spectral shape is not yet well measured due to the low statistics of the data. This analysis probes the gap between 1 and 10 PeV by using high-energy downgoing muon neutrinos. Here, to reject the large atmospheric muon background, two complementary techniques are combined. The first technique selects events with high stochasticity to reject atmospheric muon bundles whose stochastic energy losses are smoothed due to high muon multiplicity. The second technique vetoes atmospheric muons with the IceTop surface array. Using 9 yrs of data, we found two neutrino candidate events in the signal region, consistent with expectation from background, each with relatively high signal probabilities. A joint maximum likelihood estimation is performed using this sample and an independent 9.5-yr sample of tracks to measure the neutrino spectrum. A likelihood ratio test is done to compare the single power-law (SPL) vs SPL+cutoff hypothesis; the SPL+cutoff model is not significantly better than the SPL. High-energy astrophysical objects from four source catalogs are also checked around the direction of the two events. No significant coincidence was found.

Abbasi, R. [Loyola University Chicago] (ORCID:0000↗

PISCES two-detector covariance matrix fit for the NOvA Experiment

NOvA is a long-baseline neutrino oscillation experiment with two functionally identical detectors: a Near Detector (ND) at Fermilab, placed 1 km from the neutrino source, and a Far Detector (FD) located 810 km away from the ND in Minnesota. NOvA's primary physics goals are the precise measurements of neutrino oscillation parameters $\theta_{23}$ and $\Delta m^2_{32}$ , determine the neutrino mass ordering, and constrain the value of $\delta_{CP}$, via the study of muon neutrino to electron neutrino oscillation. In the standard NOvA three-flavor analysis, oscillation parameters are extracted using an extrapolation technique in which the ND data constrain the FD prediction through a ratio method. While this allows for systematic uncertainties sharing the same effects in both detectors to cancel, it remains an FD-only fit and does not fully leverage the constraining power of the high-statistics ND. This analysis proposes a simultaneous ND+FD fit using the PISCES method. PISCES (Parameter Inference with Systematic Covariance and Exact Statistics) is a framework designed to support complex configurations such as a joint ND+FD fit. This allows PISCES to take full advantage of the ND data to directly constrain systematic uncertainties across all samples. In PISCES, systematic uncertainties are encoded in a fractional covariance matrix, and statistical uncertainties are handled with a Poisson likelihood, making the approach well suited for low-statistics samples. For interpretability, we further use a Newton–Raphson + PCA method to recover per-systematic pulls from the covariance formulation. This poster presents the full PISCES joint ND+FD fit for the NOvA three-flavor analysis, describes its implementation and evaluates its performance through extensive robustness tests and fake data studies. It also provides a comparison between the PISCES joint ND+FD results and the standard NOvA extrapolation method.

Rajaoalisoa, Miriama [Cincinnati U.] (ORCID:000000↗

Stacked reverberation mapping of high-redshift quasars in DESI. I. Feasibility analysis

The broad-line region of quasars has long been probed by reverberation mapping techniques that measure time lags between continuum and broad emission-line variations. Stacked reverberation mapping has been proposed as a less observationally expensive alternative to traditional methods. This ensemble approach also reduces biases from small-number statistics. The Dark Energy Spectroscopic Instrument (DESI) is conducting the most extensive spectroscopic survey of quasars to date. We create mock light curves emulating expected DESI quasar observations at redshifts $1.48\lt z\lt 5.2$ and luminosities $44.68 \le \log \lambda L_{1350 \mathring{\rm A}{}} / \mathrm{erg\, s^{-1}} \le 45.99$ to test stacked reverberation mapping feasibility using sparse spectroscopic data paired with well-sampled photometric data. The pipeline, using the lag estimation code JAVELIN (Just Another Vehicle for Estimating Lags In Nuclei), successfully recovers the simulated C IV lags within 1σ of the true values using spectroscopic light curves composed of only a few spectral epochs (2–10) with irregular cadences. We investigate how observational factors, including C IV flux error magnitude, number of stacked quasars, and spectral epoch count, affect performance. This work motivates a pathway for future stacked reverberation mapping projects with large-scale spectroscopic surveys of quasars having $\ge 2$ spectroscopic observations. Our results suggest an economical alternative for constraining and extending the radius–luminosity relation to higher redshifts and luminosities. Subsequently, this relation can be employed more reliably in single-epoch black hole mass measurements and quasar cosmology in these distant regimes.

quasars: general, quasars: supermassive black hole↗

Automatic detection of low surface brightness galaxies from Sloan Digital Sky Survey images

ABSTRACT Low surface brightness (LSB) galaxies are galaxies with central surface brightness fainter than the night sky. Due to the faint nature of LSB galaxies and the comparable sky background, it is difficult to search LSB galaxies automatically and efficiently from large sky survey. In this study, we established the low surface brightness galaxies autodetect (LSBG-AD) model, which is a data-driven model for end-to-end detection of LSB galaxies from Sloan Digital Sky Survey (SDSS) images. Object-detection techniques based on deep learning are applied to the SDSS field images to identify LSB galaxies and estimate their coordinates at the same time. Applying LSBG-AD to 1120 SDSS images, we detected 1197 LSB galaxy candidates, of which 1081 samples are already known and 116 samples are newly found candidates. The B-band central surface brightness of the candidates searched by the model ranges from 22 to 24 mag arcsec−2, quite consistent with the surface brightness distribution of the standard sample. A total of 96.46 per cent of LSB galaxy candidates have an axial ratio (b/a) greater than 0.3, and 92.04 per cent of them have $fracDev\_r$ < 0.4, which is also consistent with the standard sample. The results show that the LSBG-AD model learns the features of LSB galaxies of the training samples well, and can be used to search LSB galaxies without using photometric parameters. Next, this method will be used to develop efficient algorithms to detect LSB galaxies from massive images of the next-generation observatories.

79 ASTRONOMY AND ASTROPHYSICS↗

Statistical fracture behavior of doped UO 2 using a ball-on-ring equibiaxial flexure test method

Metal oxide dopants, such as titanium and chromium oxides, have garnered considerable attention for their potential to increase grain size (≥ 30 µm) in UO 2 fuel, purportedly enhancing fission gas retention during reactor operation. Fuel performance is significantly impacted by fuel fracture behavior, so it is important to understand the effects of enhanced grain size and dopant content on UO 2 fuel fracture. UO 2 pellets were doped with 0.1 wt% TiO 2 and 0.3 wt% Cr 2 O 3 to alter density and grain size. Inductively coupled plasma mass spectroscopy measured dopant levels pre- and post-sintering. X-ray diffraction revealed lattice changes and microstrain via Rietveld refinement. Field emission scanning electron microscopy determined grain sizes of approximately 30 µm for TiO 2 doping and 7 µm for Cr 2 O 3 doping. Transverse rupture strength tests were performed on over 30 samples per dataset to obtain characteristic strength and Weibull modulus. Results indicate no statistical difference in fracture strength between 0.1 wt% TiO 2 doped UO 2 and undoped UO 2 , while 0.3 wt% Cr 2 O 3 doped UO 2 exhibited a 20% decrease in fracture strength. Doped UO 2 samples also showed reduced Weibull modulus compared to undoped UO 2 , suggesting increased scatter in fracture strength. This study's findings suggest that titanium and chromium oxide doping in UO 2 , regardless of grain size, induce residual stresses, decreasing fracture strength and increasing variability in fracture behavior.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Constraining the nuclear spin distribution using improved 197 Au neutron resonance parameters

New neutron transmission data at resonance energies using a 197 Au sample were measured using an early version of the Device for Indirect Capture Experiments on Radionuclides (DICER), which is under development at the Los Alamos Neutron Science Center (LANSCE). These data were combined with previous neutron transmission and capture data in a simultaneous R-matrix analysis to extract improved neutron resonance parameters for this nuclide. As a result, total radiation widths, Γ γ , were obtained for 33 J=1 and 44 J=2 197 Au+n resonances. Γ γ distributions for these two spins states were compared to distributions calculated according to the nuclear statistical model using published nuclear level density (NLD) and photon strength functions (PSF) measured using the Oslo technique. The calculated distributions were found to be narrower and the average values for the two spins states closer together than the data. The calculation can be brought into agreement with the data by substantial modifications to the spin distribution in 198 Au as a function of excitation energy. As far as we know, the spin distribution currently is otherwise poorly constrained. The modified spin distribution changes the shapes of the NLD and PSF extracted using the Oslo technique and so could have broad implications.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

LENS: Learning Enabled Network Synthesis

RTRC and UMD have developed novel machine learning based methods under the ARPA-E DIFFERENTIATE program for rapid acceleration of hypothesis generation in complex architecture design spaces involving both discrete choices of component inclusion and interconnection and continuous parametric decisions. The project named Learning Enabled Network Synthesis (LENS) further demonstrated the developed methods on challenging electrical power converter design problems by identifying the most suitable circuit topologies and simultaneously selecting the most appropriate components to achieve optimized design of power converter with improved performances. We demonstrated that LENS could enable exploration of very large design space of circuit topologies and components by addressing the limitations of conventional design process in non-linear, high switching speed, multi-dimensional power converter design and optimization. The key innovation developed in LENS is the seamless integration of statistical learning and logical reasoning techniques and building on the individual strengths of these techniques for rapid hypothesis discovery. The main component of LENS comprises of: 1) Graph Reasoning Engine (GRE) to enforce composition rules that rapidly reject all discrete architectures that are composed incorrectly and generates an adaptive database of feasible designs which can be used by ML modules, 2) Graph Generative Learning module which is a deep neural network based generative model for graph architectures which can enable design space exploration beyond the dataset generated by the GRE, 3) Graph Reduced Order Model (ROM) for graph domains for accelerating computation of output metrics, and 4) Active learning and Rule Discovery module for sample efficient learning and extracting logical rules from the learned ML models which will be integrated in the GRE to enhance the filtering effectiveness. LENS approach can be applied to any design domains where designs can be represented as multi-attribute graphs. The LENS team integrated the various technical innovations listed above into an optimization pipeline and exercised the optimization pipeline on the converter design problem. The LENS project demonstrated that the developed AI/ML technologies can be used to generate novel converter circuits >45x faster than experts on chosen use-cases. This can enable faster design space exploration and identification of new designs which are not considered by experts due to the increasing design space complexity. This has significant potential impact on the public and energy needs of the country. It is currently estimated that 30% of all electrical powers generated passes through power converters. The future estimate is that 80% of all power generated would be passing through converters. LENS fills a critical gap in this space since by accelerating the design process the designers would be able to generate more efficient converters which can lead to significant energy savings for the country.

42 ENGINEERING↗

Mass Detection for Heavy-Duty Vehicles using Gaussian Belief Propagation

Predicting vehicle mass is critical to accurately estimate energy use and emissions of commercial trucks. However, data from vehicle telematics is often not at sufficient temporal resolution or accuracy for use in model-based detection methods. In this work, a new statistical mass prediction technique is described for heavy-duty vehicles that incorporates the use Gaussian Belief Propagation (GBP) for probabilistic inference. Similar to Bayesian inference models, the GBP model typically requires less labeled training data than other contemporary machine learning techniques. First, a factor graph is constructed, and a set of Gaussian belief nodes with associated means and variances are fitted to the training data. To better handle noisy input data, the GBP mass prediction model utilizes a k-nearest factors (kNF) algorithm for probabilistic inference on unseen testing data. The proposed method is compared with a classical weighted k-nearest neighbors (kNN) regressor. This statistical kNF-GBP model works even with low-quantity, low-quality initial training data, while being capable of realtime mass estimation. Unlike the kNN regressor, the GBP model produces a measure of uncertainty with its predictions. The proposed method is validated using curve-sampled driving data collected from multiple cloud-connected Class 8 regional haul diesel trucks. Both the kNN regressor and the kNF-GBP mass prediction model were able to predict payload mass with coefficients of determination above 0.97 with minimal data preprocessing.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗