Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian process fitting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Gaussian process guide for signal regression in magnetic fusion

Extracting reliable information from diagnostic data in tokamaks is critical for understanding, analyzing, and controlling the behavior of fusion plasmas and validating models describing that behavior. Recent interest within the fusion community has focused on the use of principled statistical methods, such as Gaussian process regression (GPR), to attempt to develop sharper, more reliable, and more rigorous tools for examining the complex observed behavior in these systems. While GPR is an enormously powerful tool, there is also the danger of drawing fragile, or inconsistent conclusions from naive GPR fits that are not driven by principled treatments. Here we review the fundamental concepts underlying GPR in a way that may be useful for broad-ranging applications in fusion science. We also revisit how GPR is developed for profile fitting in tokamaks. We examine various extensions and targeted modifications applicable to experimental observations in the edge of the DIII-D tokamak. Finally, we discuss best practices for applying GPR to fusion data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Assessing equation of state-independent relations for neutron stars with nonparametric models

Relations between neutron star properties that do not depend on the nuclear equation of state offer insights on neutron star physics and have practical applications in data analysis. Such relations are obtained by fitting to a range of phenomenological or nuclear physics equation of state models, each of which may have varying degrees of accuracy. In this study we revisit commonly used relations and reassess them with a very flexible set of phenomenological nonparametric equation of state models that are based on Gaussian processes. Our models correspond to two sets: equations of state which mimic hadronic models, and equations of state with rapidly changing behavior that resemble phase transitions. Here we quantify the accuracy of relations under both sets and discuss their applicability with respect to expected upcoming statistical uncertainties of astrophysical observations. We further propose a goodness-of-fit metric which provides an estimate for the systematic error introduced by using the relation to model a certain equation-of-state set. Overall, the nonparametric distribution is more poorly fit with existing relations, with the I–Love–Q relations retaining the highest degree of universality. Fits degrade for relations involving the tidal deformability, such as the binary-Love and compactness-Love relations, and when introducing phase transition phenomenology. For most relations, systematic errors are comparable to current statistical uncertainties under the nonparametric equation of state distributions.

79 ASTRONOMY AND ASTROPHYSICS↗

Dynamics of bottlebrush polymers

Bottlebrushes are an interesting class of polymers which shows intriguing material properties often associated with dynamics. While dynamical phenomena in linear polymers are well understood and existing theories can describe them in a good way, bottlebrush dynamics have only rarely been investigated. Therefore, we performed dielectric spectroscopy and quasi-elastic neutron scattering to study the dynamics of polydimethylsiloxane-based bottlebrush polymers, PDMS-g-PDMS focusing mostly on the segmental dynamics of the side chains. Comparing the relaxation times of the α – relaxation, tracked with dielectric spectroscopy, of bottlebrush polymers with those of their respective linear side chains show a slowing down once the side chains are attached to the backbone. This effect diminishes and finally vanishes with increasing side chain length. The time and length scale, offered by quasi-elastic neutron scattering, fits for the segmental dynamics together with faster processes. The Q -dependence of the segmental relaxation times allows to classify bottlebrush polymers as heterogenous including a non-Gaussian character. For such a dynamical system, the mean square displacement needs to be separated into single processes before an overall mean square displacement can be generated by applying the time temperature superposition principle.

36 MATERIALS SCIENCE↗

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING↗

O'Hare Airport roadway traffic prediction via data fusion and Gaussian process regression

This study proposes an approach of leveraging information gathered from multiple traffic data sources at different resolutions to obtain approximate inference on the traffic distribution of Chicago's O'Hare Airport area. Specifically, it proposes the ingestion of traffic datasets at different resolutions to build spatiotemporal models for predicting the distribution of traffic volume on the road network. Due to its good adaptability and flexibility for spatiotemporal data, the Gaussian process (GP) regression was employed to provide short-term forecasts using data collected by loop detectors (sensors) and supplemented by telematics data. The GP regression is used to make predictions of the distribution of the proportion of sensor data traffic volume represented by the telematics data for each location of the sensors. Consequently, the fitted GP model can be used to determine the approximate traffic distribution for a testing location outside of the training points. Policymakers in the transportation sector can find the results of this work helpful for making informed decisions relating to current and future transportation conditions in the area.

42 ENGINEERING↗

Predicting molecular dipole moments by combining atomic partial charges and atomic dipoles

The molecular dipole moment ( μ ) is a central quantity in chemistry. It is essential in predicting infrared and sum-frequency generation spectra as well as induction and long-range electrostatic interactions. Furthermore, it can be extracted directly—via the ground state electron density—from high-level quantum mechanical calculations, making it an ideal target for machine learning (ML). Here, we choose to represent this quantity with a physically inspired ML model that captures two distinct physical effects: local atomic polarization is captured within the symmetry-adapted Gaussian process regression framework which assigns a (vector) dipole moment to each atom, while the movement of charge across the entire molecule is captured by assigning a partial (scalar) charge to each atom. The resulting “MuML” models are fitted together to reproduce molecular μ computed using high-level coupled-cluster theory and density functional theory (DFT) on the QM7b dataset, achieving more accurate results due to the physics-based combination of these complementary terms. The combined model shows excellent transferability when applied to a showcase dataset of larger and more complex molecules, approaching the accuracy of DFT at a small fraction of the computational cost. We also demonstrate that the uncertainty in the predictions can be estimated reliably using a calibrated committee model. The ultimate performance of the models—and the optimal weighting of their combination—depends, however, on the details of the system at hand, with the scalar model being clearly superior when describing large molecules whose dipole is almost entirely generated by charge separation. These observations point to the importance of simultaneously accounting for the local and non-local effects that contribute to μ ; furthermore, they define a challenging task to benchmark future models, particularly those aimed at the description of condensed phases.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Jet fragmentation transverse momentum distributions in pp and p-Pb collisions at $ \sqrt{s}$, $\sqrt{s_{\mathrm{NN}}}$ = 5.02 TeV

Jet fragmentation transverse momentum ( j T ) distributions are measured in proton-proton (pp) and proton-lead (p-Pb) collisions at s NN = 5 . 02 TeV with the ALICE experiment at the LHC. Jets are reconstructed with the ALICE tracking detectors and electromagnetic calorimeter using the anti- k T algorithm with resolution parameter R = 0 . 4 in the pseudorapidity range |η| < 0 . 25. The j T values are calculated for charged particles inside a fixed cone with a radius R = 0 . 4 around the reconstructed jet axis. The measured j T distributions are compared with a variety of parton-shower models. Herwig and P ythia 8 based models describe the data well for the higher j T region, while they underestimate the lower j T region. The j T distributions are further characterised by fitting them with a function composed of an inverse gamma function for higher j T values (called the “wide component”), related to the perturbative component of the fragmentation process, and with a Gaussian for lower j T values (called the “narrow component”), predominantly connected to the hadronisation process. The width of the Gaussian has only a weak dependence on jet transverse momentum, while that of the inverse gamma function increases with increasing jet transverse momentum. For the narrow component, the measured trends are successfully described by all models except for Herwig. For the wide component, Herwig and PYTHIA 8 based models slightly underestimate the data for the higher jet transverse momentum region. These measurements set constraints on models of jet fragmentation and hadronisation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Chocolate Chip Cookie Model: Dust Geometry of Milky Way–like Disk Galaxies

We present a new two-component dust geometry model, the Chocolate Chip Cookie model, where the clumpy nebular regions are embedded in a diffuse stellar/interstellar medium disk, like chocolate chips in a cookie. By approximating the binomial distribution of the clumpy nebular regions with a continuous Gaussian distribution and omitting the dust scattering effect, our model solves the dust attenuation process for both the emission lines and stellar continua via analytical approaches. Our Chocolate Chip Cookie model successfully fits the inclination dependence of both the effective dust reddening of the stellar components derived from stellar population synthesis and that of the emission lines characterized by the Balmer decrement for a large sample of Milky Way–like (MW-like) disk galaxies selected from the main galaxy sample of the Sloan Digital Sky Survey. Our model shows that the clumpy nebular disk is about 0.55 times thinner and 1.6 times larger than the stellar disk for MW-like galaxies, whereas each clumpy region has a typical optical depth of τ cl,V ~ 0.5 in the V band. After considering the aperture effect, our model prediction on the inclination dependence of dust attenuation is also consistent with observations. Not only that, in our model, the dust attenuation curve of the stellar population naturally depends on the inclination, and its median case is consistent with the classical Calzetti law. As the modeling constraints are from the optical wavelengths, our model is unaffected by the optically thick dust component, which however could bias the model's prediction of the infrared emissions.

79 ASTRONOMY AND ASTROPHYSICS↗

Fermilab Booster loss modelling and rebalancing using Bayesian methods

Fermilab Booster is being upgraded for the PIP-II project to support 20Hz ramp rate at higher intensities. Loss trip limits determine the achievable peak power. To meet PIP-II requirements, losses need to be halved as compared to current levels. Losses primarily occur at injection and transition crossing, with both gradually increasing and threshold-like intensity-dependent behaviors. The existing simulation models are not yet good enough for quantitative loss predictions. In practice, it will be necessary to tune up the Booster using iterative methods and operator intuition. In this paper we present an effort to systematically model Booster losses using active learning (Bayesian exploration) techniques, and subsequently to rebalance them for higher trip limit margins. We first created several sets of spatially and temporally isolated orbit and optics knobs, and trained Gaussian process models for each beam loss monitor as well as beam current. This is a complex task due to safety and timing requirements – we discuss mitigations such as uncertainty constraints and approximate fitting. Once models are stable, we perform large-scale single and multi-objective tuning using scalarized objectives made up of critical beam loss locations. Our results demonstrate significant rebalancing of losses, increasing trip margins, as well as an overall improvement in beam transmission efficiency. We are exploring how to combine existing simulations with experimental data and automate the collection procedure so that more advanced surrogate models can be created over time.

Kuklev, Nikita [Fermilab]↗

Effective Opacity of the Intergalactic Medium from Galaxy Spectra Analysis

We measure the effective opacity (τ{sub eff}) of the intergalactic medium from the composite spectra of 281 Lyman-break galaxies in the redshift range 2 ≲ z ≲ 3. Our spectra are taken from the COSMOS Lyα Mapping And Tomographic Observations survey derived from the Low Resolution Imaging Spectrometer on the W.M. Keck I telescope. We generate composite spectra in two redshift intervals and fit them with spectral energy distribution (SED) models composed of simple stellar populations. Extrapolating these SED models into the Lyα forest, we measure the effective Lyα opacity (τ{sub eff}) in the 2.02 ≤ z ≤ 2.44 range. At z = 2.22, we estimate τ{sub eff} =0.159±0.001 from a power-law fit to the data. These measurements are consistent with estimates from quasar analyses at z < 2.5 indicating that the systematic errors associated with normalizing quasar continua are not substantial. We provide a Gaussian processes model of our results and previous τ{sub eff} measurements that describes the steep redshift evolution in τ{sub eff} from z = 1.5–4.

79 ASTRONOMY AND ASTROPHYSICS↗

BeyondPlanck: V. Minimal ADC Corrections for Planck LFI

We describe the correction procedure for Analog-to-Digital Converter (ADC) differential non-linearities (DNL) adopted in the Bayesian end-to-end BEYONDPLANCK analysis framework. This method is nearly identical to that developed for the official Planck Low Frequency Instrument (LFI) Data Processing Center (DPC) analysis, and relies on the binned rms noise profile of each detector data stream. However, rather than building the correction profile directly from the raw rms profile, we first fit a Gaussian to each significant ADC-induced rms decrement, and then derive the corresponding correction model from this smooth model. The main advantage of this approach is that only samples which are significantly affected by ADC DNLs are corrected, as opposed to the DPC approach in which the correction is applied to all samples, filtering out signals not associated with ADC DNLs. The new corrections are only applied to data for which there is a clear detection of the non-linearities, and for which they perform at least comparably with the DPC corrections. Out of a total of 88 LFI data streams (sky and reference load for each of the 44 detectors) we apply the new minimal ADC corrections in 25 cases, and maintain the DPC corrections in 8 cases. All these corrections are applied to 44 or 70 GHz channels, while, as in previous analyses, none of the 30 GHz ADCs show significant evidence of non-linearity. By comparing the BEYONDPLANCK and DPC ADC correction methods, we estimate that the residual ADC uncertainty is about two orders of magnitude below the total noise of both the 44 and 70 GHz channels, and their impact on current cosmological parameter estimation is small. However, we also show that non-idealities in the ADC corrections can generate sharp stripes in the final frequency maps, and these could be important for future joint analyses with the Planck High Frequency Instrument (HFI), Wilkinson Microwave Anisotropy Probe (WMAP), or other datasets. We therefore conclude that, although the existing corrections are adequate for LFI-based cosmological parameter analysis, further work on LFI ADC corrections is still warranted.

79 ASTRONOMY AND ASTROPHYSICS↗

A Gaussian Process Regression Reveals No Evidence for Planets Orbiting Kapteyn’s Star

Radial–velocity (RV) planet searches are often polluted by signals caused by gas motion at the star’s surface. Stellar activity can mimic or mask changes in the RVs caused by orbiting planets, resulting in false positives or missed detections. Here we use Gaussian process regression to disentangle the contradictory reports of planets versus rotation artifacts from Kapteyn’s star. To model rotation, we use joint quasiperiodic kernels for the RV and Hα signals, requiring that their periods and correlation timescales be the same. We find that the rotation period of Kapteyn’s star is 125 days, while the characteristic active-region lifetime is 694 days. Adding a planet to the RV model produces a best-fit orbital period of 100 yr, or 10 times the observing time baseline, indicating that the observed RVs are best explained by star rotation only. We also find no significant periodic signals in residual RV data sets constructed by subtracting off realizations of the best-fit rotation model and conclude that both previously reported “planets” are artifacts of the star’s rotation and activity. Our results highlight the pitfalls of using sinusoids to model quasiperiodic rotation signals.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fast Emulation of Expensive Simulations using Approximate Gaussian Processes [Slides]

Nuclear Computational Low-Energy Initiative (NUCLEI) collaboration uses Density Functional Theory (DFT) simulations to predict the structure and binding energies of nuclei over a wide range of proton (Z) and neutron (N) numbers. The DFT simulations utilize a particular parameterization of a Skyrme energy density functional called UNEDF1 which depends on 12 free parameters that must be fit to data (M Kortelainen et al 2014). Fitting involves comparing (e.g.) predicted binding energies of nuclei to experimentally measured values. We use only binding energies as observables, but DFT with UNEDF1 will predict structure (shape) observables as well. In this work, assessing the capability of approximate GP emulators to balance emulator accuracy with computational speed to facilitate improved UNEDF1 calibration. Sparse GPs are straightforward to train and accurate. Calibration is not straightforward with MCMC (using MH or HMC/NUTS). We produced reusable software for continuing and building on this work as well as accessing and using Darwin cluster compute resources

97 MATHEMATICS AND COMPUTING↗

Realistic Covariance Generation for the GPM Spacecraft

We present several different methods for generating realistic predictive covariance, including Monte-Carlo simulations and more direct linear methods which require the addition of process noise. The Monte-Carlo simulation starts with an epoch uncertainty sample basis and propagates each trial to a time of interest in the future. The variance-covariance of the state elements as well as other higher order sample statistics can be readily computed from the propagated sample. While this method preserves the nonlinear effects on the propagated uncertainty, it is computationally intensive as a statistically significant sample size must be considered in the propagation process. Moreover, if the epoch covariance is optimistic, this can result in an underestimation of the prediction error. Another method is to propagate a state sensitivity matrix simultaneously with the satellite state, which allows the epoch covariances to be propagated forward in a linear fashion. This method does not preserve the non-linearity of the satellite state uncertainty but is much less computationally intensive. The propagated state covariance is the scaled to represent the realistic level of GPM state uncertainty via a "e-tuning process." The tuning process generates an inflation factor based on the observed error statistics of the predictive satellite trajectories when compared to the definitive ones. Difference tuning strategies are considered and compared via Goodness-of-Fit method testing for the Gaussian properties of the scaled covariance.

prediction error↗

Deep Learning of Dark Energy Spectroscopic Instrument Mock Spectra to Find Damped Lyα Systems

We have updated and applied a convolutional neural network (CNN) machine-learning model to discover and characterize damped Ly α systems (DLAs) based on Dark Energy Spectroscopic Instrument (DESI) mock spectra. We have optimized the training process and constructed a CNN model that yields a DLA classification accuracy above 99% for spectra that have signal-to-noise ratios (S/N) above 5 per pixel. The classification accuracy is the rate of correct classifications. This accuracy remains above 97% for lower S/N ≈1 spectra. This CNN model provides estimations for redshift and H i column density with standard deviations of 0.002 and 0.17 dex for spectra with S/N above 3 pixel -1 . Also, this DLA finder is able to identify overlapping DLAs and sub-DLAs. Further, the impact of different DLA catalogs on the measurement of baryon acoustic oscillations (BAO) is investigated. The cosmological fitting parameter result for BAO has less than 0.61% difference compared to analysis of the mock results with perfect knowledge of DLAs. This difference is lower than the statistical error for the first year estimated from the mock spectra: above 1.7%. We also compared the performances of the CNN and Gaussian Process (GP) models. Our improved CNN model has moderately 14% higher purity and 7% higher completeness than an older version of the GP code, for S/N > 3. Both codes provide good DLA redshift estimates, but the GP produces a better column density estimate by 24% less standard deviation. A credible DLA catalog for the DESI main survey can be provided by combining these two algorithms.

79 ASTRONOMY AND ASTROPHYSICS↗

The cool-star spectral catalog: A uniform collection of IUE SWP-LOs

Over the past decade and a half of its operations, the International Ultraviolet Explorer has recorded low-dispersion spectrograms in the 1150-2000 A interval of more than 800 stars of late spectral type (F-M). The sub-2000 A region contains a number of emission lines that are key diagnostics of physical conditions in the high-excitation chromospheres and subcoronal 'transition zones' of such stars. Many of the sources have been observed a number of times, and the available collection of SWP-LO exposures in the IUE Archives exceeds 4,000. With support from the Astrophysics Data Program, we have assembled the archival material into a catalog of IUE far-UV fluxes of late-type stars. In order to ensure uniform processing of the spectra, we: (1) photometrically corrected the raw vidicon images with a custom version of the 1985 SWP ITF; (2) identified and eliminated, sharp cosmic-ray 'hits' by means of a spatial filter; (3) extracted the spectral traces with the 'optimal' (weighted-slit) strategy; and (4) calibrated them against a well-characterized reference source, the DA white dwarf G191-B2B. Our approach is similar to that adopted by the IUE Project for its 'Final Archive', but our implementation is specialized to the case of chromospheric emission-line sources. We measured the resulting SWP-LO spectra using a semi-autonomous algorithm that establishes a smooth continuum by numerical filtering, and then fits the significant emissions (or absorptions) by means of a constrained Bevington-type multiple-Gaussian procedure. The algorithm assigns errors to the fitted fluxes - or upper limits in the absence of a significant detection - according to a model based on careful measurements of the noise properties of the IUE's intensified SEC cameras. Here, we describe the 'visualization' strategies we adopted to ensure human-review of the semi-autonomous processing and measuring algorithms; the derivation of the noise model and the assignment of errors; and the structure of the final catalog as delivered to the Astrophysics Data System.

Ayres, T.↗

Reducing Ground-based Astrometric Errors with Gaia and Gaussian Processes

Stochastic field distortions caused by atmospheric turbulence are a fundamental limitation to the astrometric accuracy of ground-based imaging. This distortion field is measurable at the locations of stars with accurate positions provided by the Gaia DR2 catalog; we develop the use of Gaussian process regression (GPR) to interpolate the distortion field to arbitrary locations in each exposure. We introduce an extension to standard GPR techniques that exploits the knowledge that the 2D distortion field is curl-free. Applied to several hundred 90 s exposures from the Dark Energy Survey as a test bed, we find that the GPR correction reduces the variance of the turbulent astrometric distortions ≈12× , on average, with better performance in denser regions of the Gaia catalog. The rms per-coordinate distortion in the riz bands is typically ≈7 mas before any correction and ≈2 mas after application of the GPR model. The GPR astrometric corrections are validated by the observation that their use reduces, from 10 to 5 mas rms, the residuals to an orbit fit to riz-band observations over 5 yr of the r = 18.5 trans-Neptunian object Eris. We also propose a GPR method, not yet implemented, for simultaneously estimating the turbulence fields and the 5D stellar solutions in a stack of overlapping exposures, which should yield further turbulence reductions in future deep surveys.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗