Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bias correction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Bias Correction and Statistical Downscaling of Future Solar Irradiance Projections Using the NSRDB

Assessing renewable energy resources under future climate scenarios has been highlighted to understand potential impacts of future climate change in renewable generation on the power sector. Climate model projection has been recognized by the renewable energy community as a useful data set to analyze the impacts of future climate change on renewable resources. However, future climate projections generated from general circulation models (GCMs) contain inherent biases that need to be corrected for accurate analysis of future projections of climate variables. In addition, the coarse spatiotemporal resolution of GCMs needs to be improved for regional climate studies. In this work, we develop statistical methods to downscale future projections of global horizontal irradiance (GHI) in a computationally efficient way. Our approach builds statistical downscaling models that correct bias of climate projection of GHI and downscale the future GHI projection from daily-scale to hourly-scale. The National Solar Radiation Database (NSRDB) is used to calibrate the statistical models and validate the downscaled GHI projections across the contiguous United State (CONUS). Preliminary results show that the statistical approach efficiently downscales climate projections of GHI with a nBIAS of 3%, nMAE of 34 % and nRMSE of 46% calculated against NSRDB for CONUS. This study describes the implemented methodology and initial results as well as future research to create high-resolution climate data sets for solar energy applications.

analytical models↗

Is Bias Correction in Dynamical Downscaling Defensible?

Localized projections of 21st‐century hydroclimate variables obtained from downscaling Global Climate Model (GCM) output are central to informing regional impact assessments and infrastructure planning. Regional GCM biases can be significant and, for dynamical downscaling, can be addressed either before (a priori) or after (a posteriori) downscaling. However, a priori bias correction (APBC) has generally unexplored effects on climate change signals. Here we analyze dynamically downscaled solutions of CMIP6 GCMs over the Western U.S., with and without APBC, and quantify APBC's impact on climate change signals relative to other irreducible uncertainty sources. For temperature and precipitation, the uncertainty introduced by APBC is negligible compared to that arising from GCM choice or internal variability. Furthermore, APBC greatly reduces regional models' unrealistically high snow‐water‐equivalent (SWE) biases that result directly from GCM errors. We leverage this finding to encourage the dynamical downscaling community to adopt APBC as a standard operating procedure.

Risser, Mark D.↗

Bias Correction and Statistical Downscaling of Solar Radiation Using NA-CORDEX and the NSRDB

The current state-of-art for estimating long-term PV production uses long-term estimates of solar radiation variables, such as global horizontal irradiance (GHI), from previous years. This data is used in models such as the System Advisor Model (SAM) or PYSyst to predict annual production for a PV plant. This information is then used to estimate the production over the next 20 years (a typical plant lifetime) under the assumption that the variability over the current period is representative of the future. As the PV industry moves to extend plant lifetimes to 50 years the current assumptions of representativeness of weather may not be appropriate. This is especially true as our climate changes rapidly. To assess long-term PV production, future projections for solar radiation based on projected carbon emissions are readily available in regional and global climate models. However, climate model projections contain inherent biases that may need to be corrected for accurate analysis of future projections of climate variables. Several studies have analyzed projections of solar radiation for future years, however the accuracy of the model output compared to current and historic data has not been widely studied. Chen (2021) showed that available climate models do not accurately represent solar radiation in some cases, over-projecting GHI at the surface while under-projecting its obstructions, such as clouds and aerosols. This works aims to (1) increase understanding of the accuracy of solar radiation currently available in global and regional climate models and (2) implement bias correction through linear models based on reanalysis data compared to observed solar radiation. The latter aim will be conducted using available observed solar radiation data and modeled data from several regional climate models (RCMs). The bias correction method will be applied to projections of solar radiation resulting in a more accurate representation of the future of solar production.

climate data↗

Detect and correct bias in multi-site neuroimaging datasets

The desire to train complex machine learning algorithms and to increase the statistical power in association studies drives neuroimaging research to use ever-larger datasets. The most obvious way to increase sample size is by pooling scans from independent studies. However, simple pooling is often ill-advised as selection, measurement, and confounding biases may creep in and yield spurious correlations. In this work, we combine 35,320 magnetic resonance images of the brain from 17 studies to examine bias in neuroimaging. In the first experiment, Name That Dataset, we provide empirical evidence for the presence of bias by showing that scans can be correctly assigned to their respective dataset with 71.5% accuracy. Given such evidence, we take a closer look at confounding bias, which is often viewed as the main shortcoming in observational studies. In practice, we neither know all potential confounders nor do we have data on them. Hence, we model confounders as unknown, latent variables. Kolmogorov complexity is then used to decide whether the confounded or the causal model provides the simplest factorization of the graphical model. Finally, we present methods for dataset harmonization and study their ability to remove bias in imaging features. In particular, we propose an extension of the recently introduced ComBat algorithm to control for global variation across image features, inspired by adjusting for unknown population stratification in genetics. Overall, our results demonstrate that harmonization can reduce dataset-specific information in image features. Further, confounding bias can be reduced and even turned into a causal relationship. However, harmonization also requires caution as it can easily remove relevant subject-specific information. Code is available at https://github.com/ai-med/Dataset-Bias.

42 ENGINEERING↗

A novel approach to increase accuracy in remotely sensed evapotranspiration through basin water balance and flux tower constraints

Remote sensing-derived evapotranspiration (RSET) products capture the spatiotemporal variations of evapotranspiration (ET) from field to basin scales with unprecedented details. However, their accuracy varies across RSET estimation methods and diverse hydroclimate regions. While ET modeling efforts to account for biophysical processes and controlling parameters have made good progress in recent years, a parallel approach of integrating in-situ ET with RSET could reduce biases in RSET products. Basin water balance ET (WBET) and flux tower ET are widely applied to evaluate RSET accuracy, yet such ET measurements are rarely used for RSET bias corrections, especially for large area applications. To address this issue, we propose a novel approach: the water balance equivalence (WABE) method, which generates spatially continuous WBET for correcting biases in RSET products. The WABE method computes synthetic WBET by integrating observed WBET and flux tower-derived FLUXCOM ET, which fills the spatial gaps of observed WBET and generates a spatially continuous WBET dataset. Synthetic WBET (2002–2015 annual average) of eight-digit hydrologic unit code (HUC8) basins across the conterminous United States (CONUS), constituting 44 % (887 out of 2035 basins) of CONUS basins, was determined within 2.0 % (RMSE = 12 %) of observed WBET at CONUS and between 1–12 % (RMSE = 3–33 %) across 18 regions in CONUS. With WABE-based bias corrections, the overall annual bias of RSET decreased from 10 % (RMSE = 34 %) to 6 % (RMSE = 26 %) across 37 flux tower sites. The WABE method offers a new approach for RSET accuracy improvement and shows great promise for large area implementations with a potential to yield substantial benefits for building accurate basin water budgets and water management decisions.

Khand, Kul↗

Integrating State Data Assimilation and Innovative Model Parameterization Reduces Simulated Carbon Uptake in the Arctic and Boreal Region

Model representation of carbon uptake and storage is essential for accurate projection of the response of the arctic-boreal zone to a rapidly changing climate. Land model estimates of LAI and aboveground biomass that can have a marked influence on model projections of carbon uptake and storage vary substantially in the arctic and boreal zone, making it challenging to correctly evaluate model estimates of Gross Primary Productivity (GPP). To understand and correct bias of LAI and aboveground biomass in the Community Land Model (CLM), we assimilated the 8-day Moderate Resolution Imaging Spectroradiometer (MODIS) LAI observation and a machine learning product of annual aboveground biomass into CLM using an Ensemble Adjustment Kalman Filter (EAKF) in an experimental region including Alaska and Western Canada. Assimilating LAI and aboveground biomass reduced these model estimates by 58% and 72%, respectively. The change of aboveground biomass was consistent with independent estimates of canopy top height at both regional and site levels. The International Land Model Benchmarking system assessment showed that data assimilation significantly improved CLM's performance in simulating the carbon and hydrological cycles, as well as in representing the functional relationships between LAI and other variables. Here, to further reduce the remaining bias in GPP after LAI bias correction, we re-parameterized CLM to account for low temperature suppression of photosynthesis. The LAI bias corrected model that included the new parameterization showed the best agreement with model benchmarks. Combining data assimilation with model parameterization provides a useful framework to assess photosynthetic processes in LSMs.

58 GEOSCIENCES↗

Learning to Correct Climate Projection Biases

The fidelity of climate projections is often undermined by biases in climate models due to their simplification or misrepresentation of unresolved climate processes. While various bias correction methods have been developed to post-process model outputs to match observations, existing approaches usually focus on limited, low-order statistics, or break either the spatiotemporal consistency of the target variable, or its dependency upon model resolved dynamics. We develop a Regularized Adversarial Domain Adaptation (RADA) methodology to overcome these deficiencies, and enhance efficient identification and correction of climate model biases. Instead of pre-assuming the spatiotemporal characteristics of model biases, we apply discriminative neural networks to distinguish historical climate simulation samples and observation samples. The evidences based on which the discriminative neural networks make distinctions are applied to train the domain adaptation neural networks to bias correct climate simulations. We regularize the domain adaptation neural networks using cycle-consistent statistical and dynamical constraints. An application to daily precipitation projection over the contiguous United States shows that our methodology can correct all the considered moments of daily precipitation at approximately $1^\circ$ resolution, ensures spatiotemporal consistency and inter-field correlations, and can discriminate between different dynamical conditions. Our methodology offers a powerful tool for disentangling model parameterization biases from their interactions with the chaotic evolution of climate dynamics, opening a novel avenue toward big-data enhanced climate predictions.

58 GEOSCIENCES↗

An Alternative Ensemble Streamflow Prediction Approach Using Improved Subseasonal Precipitation Forecasts from the North America Multi-Model Ensemble Phase II

In this article, streamflow forecasting at a subseasonal time scale (10–30 days into the future) is important for various human activities. The ensemble streamflow prediction (ESP) is a widely applied technique for subseasonal streamflow forecasting. However, ESP’s reliance on the randomly resampled historical precipitation limits its predictive capability. Available dynamical subseasonal precipitation forecasts provide an alternative to the randomly resampled precipitation in ESP. Prior studies found the predictive performance of raw subseasonal precipitation forecast is limited in many regions such as the central south of the United States, which raises questions about its effectiveness in assisting streamflow forecasting. To further assess the hydrologic applicability of dynamical subseasonal precipitation forecasts, we test the subseasonal precipitation forecast from North America Multi-Model Ensemble Phase II (NMME-2) at four watersheds in the central south region of the United States. The subseasonal precipitation forecasts are postprocessed with bias correction and spatial disaggregation (BCSD) to correct bias and improve spatial resolution before replacing the randomly resampled precipitation in ESP for streamflow predictions. The performance of the resulting streamflow predictions is benchmarked with ESP. Evaluation is conducted using Kling–Gupta Efficiency (KGE), continuous ranked probability score (CRPS), probability of detection (POD), false alarm ratios (FARs), as well as reliability diagrams. Our results suggest that BCSD-corrected subseasonal precipitation forecasts lead to overall improved streamflow predictions due to added skills in winter and spring. Our results also suggest that BCSD-corrected subseasonal precipitation forecasts lead to improved predictions on the occurrence of high-percentile streamflow values above 75%. Overall, BCSD-corrected subseasonal precipitation has shown promising performance, highlighting its potential broader applications for river and flood forecasting.

54 ENVIRONMENTAL SCIENCES↗

IM3/HyperFACETS Thermodynamic Global Warming (TGW) Simulation Datasets

Publication For a thorough description of the methods, see the peer-reviewed paper: Jones, A.D., Rastogi, D., Vahmani, P. et al. Continental United States climate projections based on thermodynamic modification of historical weather. Sci Data 10, 664 (2023). https://doi.org/10.1038/s41597-023-02485-5 Overview The IM3 / HyperFACETS climate simulations provide 40-year historical (1980-2019) as well as four 80-year future simulations (2020-2099) over the U.S. The future simulations are split into near (2020-2059) and far future (2060-2099) segments. The future scenarios span a range of plausible changes in future climate (both Global Circulation Model (GCM) and Representative Concentration Pathways/Shared Socioeconomic Pathway (RCP/SSP) dimensions). The simulations provide climate variables with high spatiotemporal resolution (25 hourly variables and 207 3-hourly variables at 12 km2). The datasets are generated using dynamical downscaling with the WRF (Weather Research and Forecasting) model (version 4.2.1) and therefore preserve physical consistency across variables. WRF is a state-of-the-art, fully compressible, non-hydrostatic, mesoscale numerical weather prediction model. WRF is coupled with an urban canopy model (UCM), which resolves urban surfaces. The future scenarios were developed using a thermodynamic global warming approach where past events are replayed under a range of future warming conditions. These scenarios therefore provide a perspective on potential increases in extreme event intensity, geographic scope, and duration, with previously non-extreme conditions potentially crossing new thresholds to be considered extreme by today's standards. This approach is not intended to estimate future changes in extreme event frequency that might result from changes in large-scale atmospheric dynamics. This dataset has NOT been bias corrected. A bias corrected version of selected variables is under development and will be released here when available. Scenarios Files Data for each scenario is provided in weekly NetCDF files. 25 variables are available at hourly resolution, and 207 variables are available at three-hourly resolution. Spatial resolution is 12km and spans the conterminous United States (CONUS), including some areas of Canada and Mexico, resulting in a grid of 424 by 299 cells. The spatial projection is a Lambert Conformal Conic with the following proj-string: "+proj=lcc +lat_0=40.0000076293945 +lon_0=-97 +lat_1=30 +lat_2=45 +x_0=0 +y_0=0 +R=6370000 +units=m +no_defs". The available scenarios and simulation periods are listed below: historical | 1980 - 2019 rcp45cooler | 2020 - 2059 rcp45cooler | 2060 - 2099 rcp45hotter | 2020 - 2059 rcp45hotter | 2060 - 2099 rcp85cooler | 2020 - 2059 rcp85cooler | 2060 - 2099 rcp85hotter | 2020 - 2059 rcp85hotter | 2060 - 2099 * The first year (1979, 2019, and 2059) of data within each scenario represents a model warmup period and should not be used. These are located in the `spinup_files` directory. Historical year 2020 is considered an extra year of data beyond the simulation period and can be found in the `additional_files` directory. For information on specific variables and a more in-depth discussion of methodology, please refer to the data landing page at https://tgw-data.msdlive.org. Delta Warming Files The global and CONUS warming deltas for each scenario are provided in degrees Celsius annually and monthly. Restart Files Yearly restart files are provided for each scenario which can be used to restart the WRF model at a particular point in time. Spinup Files The first year of data within each simulation period represents a model warmup period and should not be used. The files are provided here for the sake of reproducibility. Additional Files Additional years of data are provided as an extension of the historic simulation.

Jones, Andrew D.↗

Comparing the DES-SN5YR and Pantheon+ SN cosmology analyses: investigation based on ‘evolving dark energy or supernovae systematics’?

Recent cosmological analyses measuring distances of type Ia supernovae (SNe Ia) and baryon acoustic oscillations (BAO) have all given similar hints at time-evolving dark energy. To examine whether underestimated SN Ia systematics might be driving these results, Efstathiou (2025) compared overlapping SN events between Pantheon+ and DES-SN5YR (20 per cent SNe are in common), and reported evidence for an $\sim$0.04 mag offset between the low- and high-redshift distance measurements of this subsample of events. If this offset is arbitrarily subtracted from the entire DES-SN5YR sample, the preference for evolving dark energy is reduced. In this paper, we show that this offset is mostly due to different corrections for Malmquist bias between the two samples; therefore, an object-to-object comparison can be misleading. Malmquist bias corrections differ between the two analyses for several reasons. First, DES-SN5YR used an improved model of SN Ia luminosity scatter compared to Pantheon+ but the associated scatter-model uncertainties are included in the error budget. Secondly, improvements in host mass estimates in DES-SN5YR also affected SN standardized magnitudes and their bias corrections. Thirdly, and most importantly, the selection functions of the two compilations are significantly different, hence the inferred Malmquist bias corrections. Even if the original scatter model and host properties from Pantheon+ are used instead, the evidence for evolving dark energy from CMB, DESI BAO Year 1 and DES-SN5YR is only reduced from 3.9$\sigma$ to 3.3$\sigma$, consistent with the error budget. Finally, in this investigation, we identify an underestimated systematic uncertainty related to host galaxy property uncertainties, which could increase the final DES-SN5YR error budget by 3 per cent. In conclusion, we confirm the validity of the published DES-SN5YR results.

79 ASTRONOMY AND ASTROPHYSICS↗

High-Resolution WRF-Based Downscaling of Earth System Model Projections for Energy Applications across CONUS

Evaluating energy resources under future scenarios requires meteorological information that adequately resolves regional-scale variability and is suitable for regional energy system studies. Although Earth system model (ESM) outputs provide essential large-scale context, their coarse resolution and inherent biases limit direct use in energy system applications. This work presents a high-resolution dynamical downscaling framework using the Weather Research and Forecasting (WRF) model to generate energy-relevant regional fields for future scenarios over the contiguous United States (CONUS). The framework first identifies an optimal WRF configuration through sensitivity experiments, then evaluates raw and bias-corrected ESM initial and boundary conditions, with soil moisture and soil temperature bias correction implemented as an integral component of the bias-corrected ESM atmospheric forcing prior to WRF dynamical downscaling to improve land-atmosphere interactions. Simulations performed at 4-km resolution show that uncorrected ESM forcing leads to systematically dry and cold soil states, which propagate into elevated near-surface air temperature and solar irradiance biases, particularly during summer for the period 2000-2014. Incorporating bias-corrected atmospheric forcing together with soil state bias correction substantially reduces these errors and improves the representation of surface energy processes in WRF simulations. The results highlight the importance of bias-aware initialization strategies in high-resolution dynamical downscaling for future energy system analysis and planning.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-Resolution ESM Projections for Energy Applications Over the CONUS

Assessing energy resources under future scenarios requires high-resolution meteorological information that is physically consistent and suitable for regional-scale analysis. While Earth system model (ESM) projections provide valuable large-scale information, their coarse resolution and systematic biases limit direct applicability for energy system modeling and planning. In this study, we develop a high-resolution dynamical downscaling framework based on the Weather Research and Forecasting (WRF) model to translate global-scale ESM data into energy-relevant regional projections over the contiguous United States (CONUS). The framework identifies an optimized WRF configuration through numerical experiments and evaluates raw and bias-corrected ESM initial and boundary conditions, with soil moisture (SM) and soil temperature (ST) bias correction implemented as an integral part of the bias-corrected ESM forcing to improve land-atmosphere coupling prior to WRF dynamical downscaling. Using an optimized WRF configuration at 4-km resolution, we show that raw ESM forcing introduces systematic dry and cold soil biases that propagate into pronounced warm biases in near-surface air temperature and positive biases in solar irradiance, particularly during summer. Applying bias-corrected atmospheric forcing together with bias-corrected SM and ST substantially reduces these downstream biases and improves the surface energy balance and near-surface atmospheric fields. These results demonstrate that bias-aware treatment of initial conditions is critical for producing high-resolution downscaled projections suitable for energy system modeling and planning applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Redshifting galaxies from DESI to JWST CEERS: Correction of biases and uncertainties in quantifying morphology

Observations of high-redshift galaxies with unprecedented detail have now been rendered possible with the James Webb Space Telescope (JWST). However, accurately quantifying their morphology remains uncertain due to potential biases and uncertainties. To address this issue, we used a sample of 1816 nearby DESI galaxies, with a stellar mass range of 10 9.75 - 11.25 M ⊙ , to compute artificial images of galaxies of the same mass located at 0.75 ≤ z ≤ 3 and observed at rest-frame optical wavelength in the Cosmic Evolution Early Release Science (CEERS) survey. We analyzed the effects of cosmological redshift on the measurements of Petrosian radius (R p ), half-light radius (R 50 ), asymmetry (A), concentration (C), axis ratio (q), and Sérsic index (n). Our results show that R p and R 50 , calculated using non-parametric methods, are slightly overestimated due to PSF smoothing, while R 50 , q, and n obtained through fitting a Sérsic model does not exhibit significant biases. By incorporating a more accurate noise effect removal procedure, we improve the computation of A over existing methods, which often overestimate, underestimate, or lead to significant scatter of noise contributions. Due to PSF asymmetry, there is a minor overestimation of A for intrinsically symmetric galaxies. However, for intrinsically asymmetric galaxies, PSF smoothing dominates and results in an underestimation of A, an effect that becomes more significant with higher intrinsic A or at lower resolutions. Moreover, PSF smoothing also leads to an underestimation of C, which is notably more pronounced in galaxies with higher intrinsic C or at lower resolutions. We developed functions based on resolution level, defined as R p /FWHM, for correcting these biases and the associated statistical uncertainties. Applying these corrections, we measured the bias-corrected morphology for the simulated CEERS images and we find that the derived quantities are in good agreement with their intrinsic values – except for A, which is robust only for angularly large galaxies where R p /FWHM ≥ 5. Our correction functions can be applied to other surveys, offering valuable tools for future studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Redshift evolution of the underlying type Ia supernova stretch distribution

The detailed nature of type Ia supernovae (SNe Ia) remains uncertain, and as survey statistics increase, the question of astrophysical systematic uncertainties arises, notably that of the evolution of SN Ia populations. We study the dependence on redshift of the SN Ia SALT2.4 light-curve stretch, which is a purely intrinsic SN property, to probe its potential redshift drift. The SN stretch has been shown to be strongly correlated with the SN environment, notably with stellar age tracers. We modeled the underlying stretch distribution as a function of redshift, using the evolution of the fraction of young and old SNe Ia as predicted using the SNfactory dataset, and assuming a constant underlying stretch distribution for each age population consisting of Gaussian mixtures. We tested our prediction against published samples that were cut to have marginal magnitude selection effects, so that any observed change is indeed astrophysical and not observational in origin. In this first study, there are indications that the underlying SN Ia stretch distribution evolves as a function of redshift, and that the age drifting model is a better description of the data than any time-constant model, including the sample-based asymmetric distributions that are often used to correct Malmquist bias at a significance higher than 5σ. The favored underlying stretch model is a bimodal one, composed of a high-stretch mode shared by both young and old environments, and a low-stretch mode that is exclusive to old environments. The precise effect of the redshift evolution of the intrinsic properties of a SN Ia population on cosmology remains to be studied. The astrophysical drift of the SN stretch distribution does affect current Malmquist bias corrections, however, and thereby the distances that are derived based on SN that are affected by observational selection effects. We highlight that this bias will increase with surveys covering increasingly larger redshift ranges, which is particularly important for the Large Synoptic Survey Telescope.

79 ASTRONOMY AND ASTROPHYSICS↗

Improved Treatment of Host-galaxy Correlations in Cosmological Analyses with Type Ia Supernovae

Improving the use of Type Ia supernovae (SNe Ia) as standard candles requires a better approach to incorporate the relationship between SNe Ia and the properties of their host galaxies. Using a spectroscopically confirmed sample of ~1600 SNe Ia, we develop the first empirical model of underlying populations for SNe Ia light-curve properties that includes their dependence on host-galaxy stellar mass; we find a significant correlation between stretch population and stellar mass (99.9% confidence) and a weaker correlation between color and stellar mass (90% confidence). These populations are important inputs to simulations that are used to model selection effects and correct distance biases within the BEAMS with Bias Correction (BBC) framework. Here we improve BBC to also account for SNe Ia-host correlations, and we validate this technique on simulated data samples. Here, we recover the input relationship between SNe Ia luminosity and host-galaxy stellar mass (the mass step, γ) with a bias of 0.004 ±0.001 mag, which is a factor of 5 improvement over previous methods that have a γ bias of ~0.02 ± 0.001 mag. We adapt BBC for a novel dust-based model of intrinsic brightness variations, which results in a greatly reduced mass step for data (γ = 0.017 ± 0.008) and for simulations (γ = 0.006 ± 0.007). Analyzing simulated SNe Ia, the biases on the dark energy equation of state, w, vary from Δw = 0.006(5) to 0.010(5) with our new BBC method; these biases are significantly smaller than the 0.02(5) w bias using previous BBC methods that ignore SNe Ia-host correlations.

79 ASTRONOMY AND ASTROPHYSICS↗

A machine learning approach for efficient multi-dimensional integration

Many physics problems involve integration in multi-dimensional space whose analytic solution is not available. The integrals can be evaluated using numerical integration methods, but it requires a large computational cost in some cases, so an efficient algorithm plays an important role in solving the physics problems. We propose a novel numerical multi-dimensional integration algorithm using machine learning (ML). After training a ML regression model to mimic a target integrand, the regression model is used to evaluate an approximation of the integral. Then, the difference between the approximation and the true answer is calculated to correct the bias in the approximation of the integral induced by ML prediction errors. Because of the bias correction, the final estimate of the integral is unbiased and has a statistically correct error estimation. Three ML models of multi-layer perceptron, gradient boosting decision tree, and Gaussian process regression algorithms are investigated. The performance of the proposed algorithm is demonstrated on six different families of integrands that typically appear in physics problems at various dimensions and integrand difficulties. The results show that, for the same total number of integrand evaluations, the new algorithm provides integral estimates with more than an order of magnitude smaller uncertainties than those of the VEGAS algorithm in most of the test cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fair Bagging Boosting Models [SWR-24-38]

Fair Bagging Boosting Models is a software implementation of a framework for building, measuring bias and correcting bias in 3 popular forest machine learning models: gradient boosted trees (GBT), random forest (RF), and XGBoost models, using the XGBoost library. The framework takes advantage of the flexibility in XGBoost library to represent gradient boosted tree and random forest models, as well as the ability to use custom loss function.

Ugirumurera, Juliette↗