Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

VIC-Global Parameter Dataset Sensitivity with the Variable Infiltration Capacity Model: Evaluating the importance of dynamic land surface parameters when using the VIC-Global parameter dataset

Accurate prediction of runoff is essential to water resources management, flood risk assessment, and ecosystem protection. However, many hydrological models still have relatively substantial limitations when representing the influence of land use and land cover (LULC) on runoff generation and routing. Changes in LULC, such as deforestation, urban expansion, agricultural intensification, and wetland loss, have been shown to alter the water balance at the land surface through fundamental hydrologic processes (e.g., interception, infiltration, evapotranspiration, and soil storage). However, it remains an open question what the exact magnitude and timing of these impacts are for the spatial and temporal scales commonly used in engineering applications. In this analysis we focus on one aspect of recent LULC change for assessing human impacts, which is urbanization. Specifically we seek to determine the impacts of urbanization on the magnitude and timing of surface runoff and baseflow in HUC-12 basins in Clark County, Nevada which has experienced rapid urbanization. We use the Variable Infiltration Capacity (VIC) hydrology model with a widely used off-the-shelf dataset of land surface parameters, VIC-Global, both of which have been commonly used in the past for water and energy balance modeling for large scale hydrologic studies. We examine two scenarios where the first scenario removes all urbanized land cover and parameterizes those areas of the basins as barren or open shrubland. The second scenario tests the opposite case where all areas of the basins are classified as urban regardless of their present classification. The results from the VIC model show there is a low sensitivity for daily surface runoff between scenarios. The daily baseflow values indicate similar low sensitivity to the classification change during specific periods, but then have substantial differences during other period when large precipitation events are occurring. This is likely due to the assumed parameter values for the urban land cover classification made by the VIC-Global dataset. Using a static land cover parameterization is reasonable for large domain hydrology models that are being used for near-term planning horizons (<30 years). However, longer planning horizons where feedbacks between the atmosphere and land surface are important, especially in transient climate situations, considerations for how to update land surface parameters should be incorporated.

42 ENGINEERING↗

Model Assumptions and Data Characteristics: Impacts on Domain Adaptation in Building Segmentation

Studies on domain adaptation (DA) for remote sensing (RS) imagery analysis lack consistency in selection and description of evaluation scenarios. Without properly characterizing datasets, model assumptions, and evaluation scenarios, it is difficult to objectively compare DA methods and reach conclusions about their suitability across different applications. With this motivation, this work seeks to empirically assess to which extent the interaction between data characteristics and model assumptions influences the effectiveness of DA methods. Using the widely explored task of building footprint segmentation as a case study, we perform a large-scale study across over 200 DA scenarios that include variations across view angles, areas observed, and sensors used for data acquisition. Rather than adopting different model architectures or optimization criteria, we contrast the performances of two DA methods based on adversarial learning that differ only in their assumptions about source and target domains. Informed by metadata and data characteristics unveiled using traditional computer vision (CV) techniques as well as pretrained deep models, we provide a detailed meta-analysis of experiments highlighting the importance of accurately considering data assumptions for DA in RS segmentation tasks. As demonstrated by a “cherry-picking” exercise, different claims regarding which model is best could be made by selecting different subsets of evaluation scenarios. While well-calibrated assumptions can be beneficial, mismatching assumptions can lead to negative biases in DA applications. Furthermore, this study intends to motivate the community toward more consistent evaluation protocols while providing recommendations and insights toward creating novel benchmark datasets, documenting data characteristics, application-specific knowledge, and model assumptions.

42 ENGINEERING↗

The HydroBio Dataset: a new data resource for evaluating existing and potential hydropower capacity and freshwater biodiversity in the conterminous United States

Hydropower is a critical source of affordable and reliable electricity and energy system stability services in the United States. Opportunities to expand US hydropower production include retrofitting existing non-powered dams to produce power, retrofitting existing hydropower dams to improve efficiency or increase capacity, or constructing new hydropower infrastructure on currently unregulated river reaches. We created the HydroBio Dataset, which summarizes existing and potential hydropower capacity and freshwater biodiversity at the sub-basin scale in the conterminous US to contextualize existing and potential grid contributions with the freshwater ecosystems in which dams are situated. We demonstrate a use-case of this dataset by rescaling and comparing potential non-powered dam nominal capacity to rarity-threat-weighted freshwater species richness for sub-basins where both types of data exist. On average, normalized freshwater biodiversity exceeded normalized potential non-powered dam nominal capacity in these sub-basins. Potential non-powered dam nominal capacity was concentrated in sub-basins in the Upper Mississippi and Ohio hydrologic regions while freshwater biodiversity was concentrated in the South Atlantic-Gulf, Ohio, and Tennessee hydrologic regions. Additionally, non-powered dams and existing hydropower dams are located in sub-basins with similar indices of freshwater biodiversity. The HydroBio Dataset adds an additional ecological dimension of context to our understanding of current and potential future US hydropower capabilities and is a valuable decision support tool for stakeholders tasked with balancing gains in services to the US power grid with the public and environmental benefits of freshwater ecosystems.

Biodiversity↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging↗

PV Module Operating Temperature - Data and Resources

The Photovoltaic Systems Evaluation Laboratory (PSEL) at Sandia National Laboratories (SNL) in Albuquerque, NM has an extensive test site where PV modules and other system components are deployed and monitored for testing and evaluation. For this dataset PV Performance Labs has assembled one year of measurements from the Systems Long-Term Evaluation (SLTE) project (formerly known as PV Lifetime) providing the main variables needed to investigate and validate PV module operating temperature models: irradiance, ambient temperature, wind speed and back-of-module temperature. For use with more advanced thermal modeling, an estimate of down-welling long-wave radiation is also included.

14 SOLAR ENERGY↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

Using pile-up collisions as an abundant source of low-energy hadronic physics processes in ATLAS and an extraction of the jet energy resolution

During the 2015–2018 data-taking period, the Large Hadron Collider delivered proton-proton bunch crossings at a centre-of-mass energy of 13 TeV to the ATLAS experiment at a rate of roughly 30 MHz, where each bunch crossing contained an average of 34 independent inelastic proton-proton collisions. The ATLAS trigger system selected roughly 1 kHz of these bunch crossings to be recorded to disk. Offline algorithms then identify one of the recorded collisions as the collision of interest for subsequent data analysis, and the remaining collisions are referred to as pile-up. Pile-up collisions represent a trigger-unbiased dataset, which is evaluated to have an integrated luminosity of 1.33 pb -1 in 2015–2018. This is small compared with the normal trigger-based ATLAS dataset, but when combined with vertex-by-vertex jet reconstruction it provides up to 50 times more dijet events than the conventional single-jet-trigger-based approach, and does so without adding any additional cost or requirements on the trigger system, readout, or storage. The pile-up dataset is validated through comparisons with a special trigger-unbiased dataset recorded by ATLAS, and its utility is demonstrated by means of a measurement of the jet energy resolution in dijet events, where the statistical uncertainty is significantly reduced for jet transverse momenta below 65 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Variation in forest root image annotation by experts, novices, and AI

Abstract Background The manual study of root dynamics using images requires huge investments of time and resources and is prone to previously poorly quantified annotator bias. Artificial intelligence (AI) image-processing tools have been successful in overcoming limitations of manual annotation in homogeneous soils, but their efficiency and accuracy is yet to be widely tested on less homogenous, non-agricultural soil profiles, e.g., that of forests, from which data on root dynamics are key to understanding the carbon cycle. Here, we quantify variance in root length measured by human annotators with varying experience levels. We evaluate the application of a convolutional neural network (CNN) model, trained on a software accessible to researchers without a machine learning background, on a heterogeneous minirhizotron image dataset taken in a multispecies, mature, deciduous temperate forest. Results Less experienced annotators consistently identified more root length than experienced annotators. Root length annotation also varied between experienced annotators. The CNN root length results were neither precise nor accurate, taking ~ 10% of the time but significantly overestimating root length compared to expert manual annotation ( p = 0.01). The CNN net root length change results were closer to manual ( p = 0.08) but there remained substantial variation. Conclusions Manual root length annotation is contingent on the individual annotator. The only accessible CNN model cannot yet produce root data of sufficient accuracy and precision for ecological applications when applied to a complex, heterogeneous forest image dataset. A continuing evaluation and development of accessible CNNs for natural ecosystems is required.

Handy, Grace↗

Investigating Scientific Data Change with User Research Methods

Scientific datasets are continually expanding and changing due to fluctuations with instruments, quality assessment and quality control processes, and modifications to software pipelines. Datasets include minimal information about these changes or their effects requiring scientists manually assess modifications through a number of labor intensive and ad-hoc steps. The Deduce project is investigating data change to develop metrics, methods, and tools that will help scientists systematically identify and make decisions around data changes. Currently, there is a lack of understanding, and common practices, for identifying and evaluating changes in datasets since systematically measuring and managing data change is under explored in scientific work. We are conducting user research to address this need by exploring scientist's conceptualizations, behaviors, needs, and motivations when dealing with changing datasets. Our user research utilizes multiple methods to produce foundational, generative insights and evaluate research products produced by our team. In this paper, we detail our user research process and outline our findings about data change that emerge from our studies. Our work illustrates how scientific software teams can push beyond just usability testing user interfaces or tools to better probe the underlying ideas they are developing solutions to address.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Dynamical Downscaling of Earth System Model Data for Energy System Analysis

Assessing energy resources (e.g., solar, wind, and hydro) under future scenarios requires datasets with sufficient spatial and temporal detail to capture variability and extreme events. While global-scale Earth System Model (ESM) projections are widely used, their coarse resolution limits direct application to regional energy system analyses. Dynamical downscaling offers a robust approach to generate physically consistent, fine-scale datasets that better represent local atmospheric processes impacting energy resources. In this work, we present a two-stage approach for producing high-resolution historical and future projections over the contiguous United States (CONUS). First, we optimize the Weather Research and Forecasting (WRF) model configuration for energy-relevant variables - solar irradiance, wind speed, and precipitation - by conducting ERA5-driven simulations at 8-km and 28-km resolution. Multiple physics schemes and model configurations within the WRF are evaluated against observational datasets including the National Solar Radiation Database (NSRDB), the Parameter-elevation Regressions on Independent Slopes Model (PRISM), and the Stage IV multi-radar/multi-sensor precipitation product for the CONUS domain. Using the best-performing configuration, we dynamically downscale MPI-ESM1-2-HR simulations for 2000-2060 under SSP2-4.5 and SSP5-8.5 scenarios at 4-km spatial and hourly temporal resolution. This presentation will provide a comprehensive analysis of the results from multiple numerical experiments and high-resolution ESM projections. In addition, we will discuss potential applications of our high-resolution datasets within the energy sector and outline future research avenues dedicated to evaluating how extreme weather events influence system performance and resilience.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

Evaluation of precipitation indices in suites of dynamically and statistically downscaled regional climate models over Florida

Abstract The present work evaluates historical precipitation and its indices defined by the Expert Team on Climate Change Detection and Indices (ETCCDI) in suites of dynamically and statistically downscaled regional climate models (RCMs) against NOAA’s Global Historical Climatology Network Daily (GHCN-Daily) dataset over Florida. The models examined here are: (1) nested RCMs involved in the North American CORDEX (NA-CORDEX) program, (2) variable resolution Community Earth System Models (VR-CESM), (3) Coupled Model Intercomparison Project phase 5 (CMIP5) models statistically downscaled using localized constructed analogs (LOCA) technique. To quantify observational uncertainty, three in situ-based (PRISM, Livneh, CPC) and three reanalysis (ERA5, MERRA2, NARR) datasets are also evaluated against the station data. The reanalyses and dynamically downscaled RCMs generally underestimate the magnitude of the monthly precipitation and the frequency of the extreme rainfall in summer. The models forced with CanESM2 miss the phase of the seasonality of extreme precipitation. All models and reanalyses severely underestimate both the mean and interannual variability of mean wet-day precipitation (SDII), consecutive dry days (CDD), and overestimate consecutive wet days (CWD). Metric analysis suggests large uncertainty across NA-CORDEX models. Both the LOCA and VR-CESM models perform better than the majority of models. Overall, RegCM4 and WRF models perform poorer than the median model performance. The performance uncertainty across models is comparable to that in the reanalyses. Specifically, NARR performs poorer than the median model performance in simulating the mean indices and MERRA2 performs worse than the majority of models in capturing the interannual variability of the indices.

54 ENVIRONMENTAL SCIENCES↗

Utah FORGE Well 16A(78)-32 Stimulation DFN Fracture Plane Evaluation and Data

This dataset includes files used to fit planar fractures through the preliminary earthquake catalogs of the three stages of the April 2022 well 16A(78)-32 stimulation which is linked bellow. These planar features have been used to update the FORGE reference Discrete Fracture Network (DFN) model. The files are provided to encourage other modelers to use additional workflows to find additional/alternative features. To this end, the dataset includes the cleaned earthquake catalog data translated to the FORGE reference model global reference frame, the well trajectory of 16A(78)-32 in those same coordinates, the fit 15 planar features in csv format, and a pdf file with slides illustrating the process used to fit the features. A recorded presentation of this material is available from the October 2022 FORGE Modeling and Simulation Forum which is also linked below.

15 GEOTHERMAL ENERGY↗

Predicting High Energy Arcing Fault Zones of Influence for Aluminum Using an Arc Flash Modeling Approach: Evaluation of a model bias, uncertainty, parameter sensitivity and zone of influence estimation

This report documents the development of an arc flash hazard model to calculate the incident energy and zone of influence from high energy arcing faults involving aluminum. The NRC has identified the potential for (HEAFs) involving aluminum to increase the damage zone beyond what is currently postulated in fire probabilistic risk assessment (PRA) methodologies. To estimate the hazard from HEAFs involving aluminum an arc flash model was developed. Differences between the initial model and nuclear power plant (NPP) fire PRA scenarios were identified. Modification of the initial model established from existing literature and test data was used to minimize these differences. The developed model was evaluated against NRC datasets to understand the model prediction and relative uncertainties. Finally, a range of fire PRA zone of influences (ZOI) were developed based on the developed model, target fragility estimates and update HEAF PRA methodology. The results were developed to support an NRC LIC-504 evaluation in tandem with other modeling efforts. The report documents the effort and provides a reference for any future advancements in arc flash modeling.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Comparative study of machine learning techniques for post-combustion carbon capture systems

Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.

97 MATHEMATICS AND COMPUTING↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗