Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “probabilistic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ODIN: Characterizing the Three-dimensional Structure of Two Protocluster Complexes at z = 3.1

We present a detailed study of the 3D morphology of two extended associations of multiple protoclusters at z = 3.1. These protocluster "complexes," designated COSMOS-z3.1-A and COSMOS-z3.1-C, are the most prominent overdensities of z = 3.1 Lyα emitters (LAEs) identified in the COSMOS field by the One-hundred-deg$^{2}$ DECam Imaging in Narrowbands survey. These protocluster complexes have been followed up with extensive spectroscopy from Keck, Gemini, and DESI. Using a probabilistic method that combines photometrically selected and spectroscopically confirmed LAEs, we reconstruct the 3D structure of these complexes on scales of ≈50 cMpc. We validate our reconstruction method using the IllustrisTNG300-1 cosmological hydrodynamical simulation and show that it consistently outperforms approaches relying solely on spectroscopic data. The resulting 3D maps reveal that both complexes are irregular and elongated along a single axis, emphasizing the impact of sightline on our perception of structure morphology. The complexes consist of multiple density peaks, 10 in COSMOS-z3.1-A and 4 in COSMOS-z3.1-C. The former is confirmed to be a proto-supercluster, similar to Hyperion at z = 2.4 but observed at an even earlier epoch. Multiple "tails" connected to the cores of the density peaks are seen, likely representing cosmic filaments feeding into these extremely overdense regions. The 3D reconstructions further provide strong evidence that Lyα blobs preferentially reside in the outskirts of the highest density regions. Descendant mass estimates of the density peaks suggest that COSMOS-z3.1-A and COSMOS-z3.1-C will evolve to become ultramassive structures by z = 0, with total masses log ( M / M ,⊙ ,) ≳ 15.3 , exceeding that of Coma.

Ramakrishnan, Vandana [Purdue U., West Lafayette] ↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods↗

A probabilistic inverse prediction method for predicting plutonium processing conditions

In the past decade, nuclear chemists and physicists have been conducting studies to investigate the signatures associated with the production of special nuclear material (SNM). In particular, these studies aim to determine how various processing parameters impact the physical, chemical, and morphological properties of the resulting special nuclear material. By better understanding how these properties relate to the processing parameters, scientists can better contribute to nuclear forensics investigations by quantifying their results and ultimately shortening the forensic timeline. This paper aims to statistically analyze and quantify the relationships that exist between the processing conditions used in these experiments and the various properties of the nuclear end-product by invoking inverse methods. In particular, these methods make use of Bayesian Adaptive Spline Surface models in conjunction with Bayesian model calibration techniques to probabilistically determine processing conditions as an inverse function of morphological characteristics. Not only does the model presented in this paper allow for providing point estimates of a sample of special nuclear material, but it also incorporates uncertainty into these predictions. This model proves sufficient for predicting processing conditions within a standard deviation of the observed processing conditions, on average, provides a solid foundation for future work in predicting processing conditions of particles of special nuclear material using only their observed morphological characteristics, and is generalizable to the field of chemometrics for applicability across different materials.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Copacabana: a probabilistic membership assignment method for galaxy clusters

Cosmological analyses using galaxy clusters in optical/near-infrared photometric surveys require robust characterization of their galaxy content. Precisely determining which galaxies belong to a cluster is crucial. In this paper, we present the COlor Probabilistic Assignment of Clusters And BAyesiaN Analysis (Copacabana) algorithm. Copacabana computes membership probabilities for all galaxies within an aperture centred on the cluster using photometric redshifts, colours, and projected radial probability density functions. We use simulations to validate Copacabana and we show that it achieves up to 89 per cent membership accuracy with a mild dependence on photometric redshift uncertainties and choice of aperture size. We find that the precision of the photometric redshifts has the largest impact on the determination of the membership probabilities followed by the choice of the cluster aperture size. We also quantify how much these uncertainties in the membership probabilities affect the stellar mass–cluster mass scaling relation, a relation that directly impacts cosmology. Using the sum of the stellar masses weighted by membership probabilities (⁠μ * ⁠) as the observable, we find that Copacabana can reach an accuracy of 0.06 dex in the measurement of the scaling relation at low redshift for a Legacy Survey of Space and Time type survey. These results indicate the potential of Copacabana and μ * to be used in cosmological analyses of optically selected clusters in the future.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep probabilistic direction prediction in 3D with applications to directional dark matter detectors

Abstract We present the first method to probabilistically predict 3D direction in a deep neural network model. The probabilistic predictions are modeled as a heteroscedastic von Mises-Fisher distribution on the sphere S 2 , giving a simple way to quantify aleatoric uncertainty. This approach generalizes the cosine distance loss which is a special case of our loss function when the uncertainty is assumed to be uniform across samples. We develop approximations required to make the likelihood function and gradient calculations stable. The method is applied to the task of predicting the 3D directions of electrons, the most complex signal in a class of experimental particle physics detectors designed to demonstrate the particle nature of dark matter and study solar neutrinos. Using simulated Monte Carlo data, the initial direction of recoiling electrons is inferred from their tortuous trajectories, as captured by the 3D detectors. For 40 keV electrons in a 70% He 30% CO 2 gas mixture at STP, the new approach achieves a mean cosine distance of 0.104 (26 ∘ ) compared to 0.556 (64 ∘ ) achieved by a non-machine learning algorithm. We show that the model is well-calibrated and accuracy can be increased further by removing samples with high predicted uncertainty. This advancement in probabilistic 3D directional learning could increase the sensitivity of directional dark matter detectors.

Computer Science↗

UNDERSTANDING THE SEMI-PROBABILISTIC APPROACHES IN STRUCTURAL RELIABILITY USED TO SET DESIGN RELIABILITY TARGETS FOR GRAPHITE COMPONENTS USING ASME BPVC METHODS

Graphite is a quasi-brittle material, resulting in random variability in tensile strength distributions. To account for the random variability in strength, HHA-3000 of the ASME BPVC provides two semi-probabilistic methods for qualifying nuclear graphite components in the design stage, the simplified and full assessments. The full and simplified assessments apply statistical methods to engineering-based design problems. This is often referred to as reliability-based design. Reliability-based design (RBD) is a method to develop reliable designs by accounting for uncertainties and result in small chances of failure when also considering safety factors. RBDs provide reliability targets using semi-probabilistic approaches. RBD is implemented in ASME BPVC HHA-3000 for nuclear graphite components, but is not specific to that application. There has been much confusion around the methods implemented in ASME BPVC HHA-3000 for qualifying nuclear graphite components. To address the confusion, this paper takes a hierarchical approach. First, the general RBD framework is presented. Then, the semi-probabilistic methods and the underlying assumptions implemented in the assessments are presented. The semi-probabilistic methods are separated from the engineering modifications that have been made to the assessments. After building the framework and underlying assumptions, the specific methods in the full and simplified assessments are explained in three steps: inputs, methods, outputs. The methods are applied to an H-451 reflector block. Tensile strength properties for other graphite grades are provided.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Stage-local partitioned two-step runge-kutta methods for large systems of ordinary differential equations

We introduce stage-local partitioned two-step Runge-Kutta methods are an extension of standard two-step Runge-Kutta methods, which are an alternative to the standard additive two-step Runge-Kutta methods currently existing in the literature. Furthermore, these new schemes are designed with an eye towards truly N-partitioned systems and leverage local stage approximations to make several computationally interesting approximations viable. Specifically, the focus on local stage approximations makes possible the construction of truly asynchronous schemes, in the parallel sense, possible. In addition, we show that an implicit-explicit approach to these schemes can lead to methods that require the inversion of only local nonlinear systems.

Applied Dynamical Systems↗

First-of-a-Kind Risk-Informed Digital Twin for Operational Decision Making

A digital twin (DT) is a digital model or a collection of models of a physical entity. DTs in the nuclear arena can be used from plant design through decommissioning. Decisions are typically a priori or made offline. Risk-informed decision making is identifying what can go wrong, its frequency, and the consequences of its failure. Ideally risk-informed decision making reflects the current state of the plant and provides a decision in real time. Traditionally, probabilistic risk assessments (PRAs) evaluate the failures of safety systems, the risk of core damage, and the offsite dose as the consequence. However, this DT evaluates the decisions on the control side rather than the protection side. It uses the same risk methods to probabilistically inform the decision-making process but in a different way. Rather than evaluating the risk of core damage, this DT evaluates the likelihood of avoiding a trip set point while maintaining plant safety. Performance-based assessments are identified via its probabilistic evaluation of operational alternatives based on system status. Because the purpose of the control system is to maintain system variables within prescribed operating ranges, upsets or challenges that can exceed a trip set point resulting in a plant transient and a challenge to plant mitigating systems based on actual plant conditions, are evaluated to safely maintain the plant within the operating ranges. The probabilistic portion of the model is autonomously and automatically adjusted, and the metric of interest (i.e. likelihood of avoiding a trip set point) is recalculated. The digital representation of the physical system (i.e. the DT) performs a deterministic performance–based assessment of the probabilistically identified alternatives identified to validate the probabilistic assessment. A decision-making algorithm selects the appropriate option based on the probabilistic and deterministic assessments and transmits a control signal to a component(s) to initiate a corrective action or informs an operator of its decision.

digital twin↗

QRF4P-NRT: Probabilistic Post-Processing of Near-Real-Time Satellite Precipitation Estimates Using Quantile Regression Forests

Accurate and reliable near-real-time satellite precipitation estimation is of great importance for operational large-scale flood forecasting and drought monitoring. The state-of-the-art precipitation post-processing model is based on a deterministic approach to construct relationships between satellites estimates and ground observations. We propose a probabilistic postprocessor, the Probabilistic Post-Processing of Near-Real-Time Satellite Precipitation Estimates using Quantile Regression Forests (QRF4P-NRT), based on quantile modeling, yielding both deterministic and probabilistic predictions. The experimental design incorporates different solutions of near-real-time predictors to further improve the model performance. Using the Integrated Multi-satellitE Retrievals Early Run for Global Precipitation Measurement Mission (IMERG-E) product as an example, we illustrate that the proposed method significantly improves the overall quality of the raw IMERG-E and is also superior to the bias-corrected product (IMERG Final Run, IMERG-F) at daily scale in a complex mountain basin. Evaluations of the corrected IMERG-E, raw IMERG-E, and IMERG-F using ground observation show that the corrected IMERG-E improves correlation coefficients (0.7), mean error (-0.14 mm/day) and root mean square error (3.3 mm/day) relative to the raw IMERG-E (0.31, -0.72 and 5.5 mm/day) and IMERG-F (0.34, -0.09 and 6.0 mm/day). The error decomposition further confirms that the QRF4P-NRT improves on the various deficiencies of the raw IMERG-E product. The ensemble assessment also demonstrates that the quantile outputs provide reliable prediction spread and sharp prediction intervals. The promising results indicate the great potential of the proposed method for probabilistic post-processing for near-real-time satellite precipitation estimates, and for further applications such as hydrological ensemble forecasting.

54 ENVIRONMENTAL SCIENCES↗

Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications

Abstract Robust quantification of predictive uncertainty is a critical addition needed for machine learning applied to weather and climate problems to improve the understanding of what is driving prediction sensitivity. Ensembles of machine learning models provide predictive uncertainty estimates in a conceptually simple way but require multiple models for training and prediction, increasing computational cost and latency. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but does not account for epistemic uncertainty. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainties with one model. This study compares the uncertainty derived from evidential neural networks to that obtained from ensembles. Through applications of the classification of winter precipitation type and regression of surface-layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. To encourage broader adoption of evidential deep learning, we have developed a new Python package, Machine Integration and Learning for Earth Systems (MILES) group Generalized Uncertainty for Earth System Science (GUESS) (MILES-GUESS) ( https://github.com/ai2es/miles-guess ), that enables users to train and evaluate both evidential and ensemble deep learning. Significance Statement This study demonstrates a new technique, evidential deep learning, for robust and computationally efficient uncertainty quantification in modeling the Earth system. The method integrates probabilistic principles into deep neural networks, enabling the estimation of both aleatoric uncertainty from noisy data and epistemic uncertainty from model limitations using a single model. Our analyses reveal how decomposing these uncertainties provides valuable insights into reliability, accuracy, and model shortcomings. We show that the approach can rival standard methods in classification and regression tasks within atmospheric science while offering practical advantages such as computational efficiency. With further advances, evidential networks have the potential to enhance risk assessment and decision-making across meteorology by improving uncertainty quantification, a longstanding challenge. This work establishes a strong foundation and motivation for the broader adoption of evidential learning, where properly quantifying uncertainties is critical yet lacking.

Schreck, John S.↗

Time domain probabilistic seismic risk analysis using ground motion prediction equations of Fourier amplitude spectra

Modeling of Fourier amplitude spectra (FAS) of seismic motions has gained much attention in engineering seismology. In the past few years, several ground motion prediction equations (GMPEs) and inter-frequency correlation structure of FAS have been established. Due to many preferable characteristics of FAS, probabilistic seismic hazard/risk analysis is rapidly changing from ergodic, spectrum acceleration Sa(T 0 )-based approach to non-ergodic, site-specific, FAS-based approach. This paper presents time domain intrusive framework for probabilistic seismic risk analysis using GMPE of FAS. Herein, methodology for time domain stochastic ground motion modeling based on GMPEs of FAS is presented in some detail. The simulated uncertain motions are modeled as a random process and represented by polynomial chaos Karhunen-Loève expansion. The random process excitations are further propagated into the uncertain structural system using Galerkin stochastic finite element method (SFEM). Probabilistic evolution of structural response is solved, and such solution is used to develop seismic risk for any damage state. The presented framework is illustrated through seismic risk analysis of a four-story building subjected to possible earthquakes from two strike slip faults. The influences of the epistemic uncertainties in source stress drop Δσ and site attenuation κ0 on seismic risk are investigated. The need for non-ergodic seismic risk analysis with source-specific and site specific characterizations is emphasized.

58 GEOSCIENCES↗

Development of a leading simulator/trailing simulator methodology as part of an integrated safety-security analysis for nuclear power plants

Nuclear power plant (NPP) risk assessment is broadly separated into disciplines of nuclear safety, security, and safeguards. Different analysis methods and computer models have been constructed to analyze each of these as separate disciplines. However, due to the complexity of NPP systems, there are risks that can span all these disciplines and require consideration of safety-security (2S) interactions which allows a more complete understanding of the relationship among these risks. In this work, a novel leading simulator/trailing simulator (LS/TS) method is introduced to integrate multiple generic safety and security computer models into a single, holistic 2S analysis. A case study is performed using this novel method to determine its effectiveness. The case study shows that the LS/TS method avoided introducing errors in simulation, compared to the same scenario performed without the LS/TS method. A second case study is then used to illustrate an integrated 2S analysis which shows that different levels of damage to vital equipment from sabotage at a NPP can affect accident evolution by several hours.

42 ENGINEERING↗

IACMI Project 4.2: Thermoplastic Composite Development for Wind Turbine Blades

(Section 5.1) Composites made from Arkema’s Elium® thermoplastic resin and Johns Manville fiberglass were researched during this project for applications in wind blade manufacturing. A techno-economic model was developed to model this wind blade manufacturing process using these materials in place of traditional composites made with thermoset resin. This model was based on manufacturing a 61.5-meter wind blade, which showed a 4.7% reduction in wind blade cost as compared traditional thermoset materials. These cost savings were not from the thermoplastic material costing less than traditional thermoset materials, but rather from decreased capital costs, faster cycle times and reduced energy requirements and labor costs. (Section 5.2) An infusion and curing model was developed for thermoplastic composite wind blades using PAM-RTM. The primary goal was to demonstrate the infusion simulation for the Elium® resin system on a 13-meter wind blade. Additionally, the exotherm temperature was predicted and compared to measurements, which showed model results within 10% of actual measurements. (Section 5.3) Composite laminate panels and composite sandwich panels with a balsa core were produced; specimens were cut and characterized. Similar composite specimens were made with Elium® thermoplastic resin and Hexion thermoset epoxy (RIMR135/RIMH1366) to enable comparisons between these resin systems. The static test methods included: tensile, compression, in-plane shear, interlaminar shear, flexural, sandwich core shear flexure, and single cantilever beam tests for sandwich beams. Fatigue testing at room temperature was completed to composite laminate panels at a stress ratio of R=0.1 and R=10. In addition, fatigue testing to laminate panels was completed at -30°C, and at room temperature after conditioning specimens at 70°C and 90% relative humidity. Overall, mechanical test results from Elium® composites are similar to epoxy composites. (Section 5.4) Elium composite panels were produced with intentional defects such as voids and nonwetting of fibers to begin to understand performance sensitivity to defects. A thermal digital image correlation (TDIC) method provides high spatial resolution strain field at elevated temperatures and can be used to identify defective regions within composite panels. Flexural modulus differences of 21% were seen between defect and non-defect panels. Other Elium® composite panels were forced to be defective by boiling the resin after infusion, which created voids throughout the composite laminate. X-ray computed tomography scanning was used to view the internal structure of the defect panels. Defect panels had a significant reduction in fatigue life as compared to baseline panels produced without intentional defects. (Section 5.5) Lap shear specimens were fabricated to compare the lap shear strength of an off-the-shelf adhesive (Plexus MA590) and two new adhesives developed by Arkema (Bostik SAF30 90 and Bostik SAF30 120). ISO standard 4587:2003 was used to standardize the testing method and sample fabrication. Lap shear specimens were made at 1mm, 3mm, and 10mm thicknesses. The Bostik adhesive lap shear test results were similar to Plexus for all thicknesses. (Section 5.6) Fiber-reinforced polymer (FRP) composites are typically used in high-performance applications (e.g., aerospace), and their expansion into high-volume industries (e.g. consumer automotive and wind turbine blade manufacturer or similar) is hindered by their cost and a lack of efficient manufacturing techniques. Monitoring the curing process of these composites during manufacturing can improve the efficiency of the process, and therefore reduce the manufacturing cost. Cure monitoring techniques were developed that use probabilistic estimation methods and surface temperature measurements made using infrared cameras. These techniques enable real-time monitoring of the infusion process to locate manufacturing flaws, and they can, potentially, estimate residual stresses in the part. Their commercialization will help facilitate expansion of FRP composites in high-volume industries. (Section 5.7) A 13-meter composite wind blade was produced with Elium® resin and Johns Manville fiberglass; this blade was made with VARTM processing similar to how megawatt-scale wind blades are currently manufactured, but no post-mold heating was used for this thermoplastic composite blade. The wind blade underwent full-scale validation for static loading (4-different load orientations) and flapwise fatigue loading to simulate 20-years of operational loads. The thermoplastic composite wind blade withstood the loading without any noted issues and performed similar to results from a previous full-scale validation to an equivalent epoxy composite wind blade produced with the same blade molds. (Section 5.8) A study was conducted to determine the feasibility of recycling composite wind turbine blade components fabricated with glass fiber reinforced Elium® thermoplastic resin. Dissolution, which is a process unique to thermoplastic matrices, allows recovery of both the polymer matrix and full-length glass fibers, while maintaining their stiffness and strength throughout the recovery process. The economics of recycling is favorable if 50% of the glass fiber is recovered and resold for a process of $\$$ 0.28/kg, and 90% of the resin is recovered and resold at a price of $\$$ 2.50/kg.(Section 10) Recommendations are outlined for commercializing thermoplastic resin for composite wind blade production, in addition to recommended areas for future research.

17 WIND ENERGY↗

Comparing Capacity Credit Calculations for Wind: A Case Study in Texas

The degree to which wind energy can contribute to the capacity needed to meet resource adequacy requirements, also known as capacity credit (CC), varies regionally with wind resource and correlation to net load. CC is an important metric widely used for resource planning and resource adequacy assessments. However, there are multiple methods for computing and estimating CC, depending on specific needs, access to data, and computational burden. It is unclear the extent to which the CC computation method may influence the result. To address this, we use a probabilistic resource adequacy tool and multiple approximation methods to systematically assess the CC of wind for near-term wind deployment under a case study in Texas. We find that proper consideration of transmission constraints is important; some approximation methods may overestimate the CC of wind due to a lack of consideration of transmission constraints, while other approximation methods may underestimate the CC by not capturing the ability of wind to be shipped to neighboring regions. In this case study, we find that several approximation methods do come close to the CC calculated by more robust probabilistic methods. However, the best approximation method may vary on a case-by-case basis, depending on system-specific considerations.

17 WIND ENERGY↗

A mathematical assessment of the isolation random forest method for anomaly detection in big data

We present the mathematical analysis of the Isolation Random Forest Method (IRF Method) for anomaly detection, proposed by Liu F.T., Ting K.M. and Zhou Z. H. in their seminal work as a heuristic method for anomaly detection in Big Data. We prove that the IRF space can be endowed with a probability induced by the Isolation Tree algorithm (iTree). In this setting, the convergence of the IRF method is proved, using the Law of Large Numbers. Here, a couple of counterexamples are presented to show that the method is inconclusive and no certificate of quality can be given, when using it as a means to detect anomalies. Hence, an alternative version of the method is proposed whose mathematical foundation is fully justified. Furthermore, a criterion for choosing the number of sampled trees needed to guarantee confidence intervals of the numerical results is presented. Finally, numerical experiments are presented to compare the performance of the classic method with the proposed one.

97 MATHEMATICS AND COMPUTING↗

An analysis of Bayesian estimates for missing higher orders in perturbative calculations

With current high precision collider data, the reliable estimation of theoretical uncertainties due to missing higher orders (MHOs) in perturbation theory has become a pressing issue for collider phenomenology. Traditionally, the size of the MHOs is estimated through scale variation, a simple but ad hoc method without probabilistic interpretation. Bayesian approaches provide a compelling alternative to estimate the size of the MHOs, but it is not clear how to interpret the perturbative scales, like the factorisation and renormalisation scales, in a Bayesian framework. Recently, it was proposed that the scales can be incorporated as hidden parameters into a Bayesian model. In this paper, we thoroughly scrutinise Bayesian approaches to MHO estimation and systematically study the performance of different models on an extensive set of high-order calculations. We extend the framework in two significant ways. First, we define a new model that allows for asymmetric probability distributions. Second, we introduce a prescription to incorporate information on perturbative scales without interpreting them as hidden model parameters. We clarify how the two scale prescriptions bias the result towards specific scale choice, and we discuss and compare different Bayesian MHO estimates among themselves and to the traditional scale variation approach. Finally, we provide a practical prescription of how existing perturbative results at the standard scale variation points can be converted to 68%/95% credibility intervals in the Bayesian approach using the new public code MiHO.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗