Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Model Calibration with Markov Chain Monte Carlo Tutorial

The purpose of this tutorial is to demonstrate how to use Markov chain Monte Carlo (MCMC) to calibrate a model. By calibration, we mean the selection of model parameters (and, when relevant, structures). A common goal in model development and diagnostics is calibration, or the identification of model structures and parameters which are consistent with data. While models can be calibrated through hand-tuning parameters or minimizing simple error metrics such as root-mean-square-error (RMSE), these approaches can underrepresent the probabilistic nature of the data-generating process, as well as the potential for multiple model configurations to be consistent with the data. Probabilistic uncertainty quantification, which is the topic of this notebook, can address these concerns. This tutorial is presented as an appendix to the e-book: Addressing Uncertainty in MultiSector Dynamics Research.

Markov chain Monte Carlo↗

Evaluation of Two Crew Module Boilerplate Tests Using Newly Developed Calibration Metrics

The paper discusses a application of multi-dimensional calibration metrics to evaluate pressure data from water drop tests of the Max Launch Abort System (MLAS) crew module boilerplate. Specifically, three metrics are discussed: 1) a metric to assess the probability of enveloping the measured data with the model, 2) a multi-dimensional orthogonality metric to assess model adequacy between test and analysis, and 3) a prediction error metric to conduct sensor placement to minimize pressure prediction errors. Data from similar (nearly repeated) capsule drop tests shows significant variability in the measured pressure responses. When compared to expected variability using model predictions, it is demonstrated that the measured variability cannot be explained by the model under the current uncertainty assumptions.

Horta, Lucas G.↗

Multiresolution Quantum Chemistry: Nonlinear Response Properties at the Basis Set Limit

We benchmark the accuracy of Dunning correlation-consistent Gaussian basis sets for computing frequencydependent second-order hyperpolarizabilities relevant to second-harmonic generation (SHG), using multiresolution analysis (MRA) as a reference. Basis set errors are analyzed using a unit-sphere representation of the effective hyperpolarizability vector, enabling direct assessment of directional error structure. We introduce a relative RMS total error metric that integrates directional deviations over the unit sphere and complement it with signed projection errors that distinguish over- and underestimation. Unsupervised clustering based on these signed directional metrics reveals four distinct convergence behaviors across a set of 68 molecules. Unitsphere visualizations of representative systems show that basis set errors are often highly anisotropic and localized along specific bond directions, even when global error measures appear small. Doubly augmented basis sets consistently outperform singly augmented ones, and core-polarization functions are required for uniform convergence in second-row systems. Overall, this work demonstrates that directional analysis combined with clustering provides a robust framework for understanding basis set convergence in nonlinear optical response properties.

Basis sets↗

Ground Validation of TRMM 3B43 V7 Precipitation Estimates Over Colombia. Part I: Monthly and Seasonal Timescales

In this study, we validate precipitation estimates remotely sensed by the Tropical Rainfall Measuring Mission (TRMM) at monthly and seasonal timescales, during the period 1998–2015, by calculating and analyzing diverse error metrics between the 3B43 V7 product and in situ measurements from 1,180 rain gauges over Colombia, of which at least 987 are fully independent of TRMM. We explore the existence of spatiotemporal patterns to assess the performance of 3B43 V7 over the five major natural regions of Colombia: Caribbean, Pacific, Andes, Orinoco and Amazon. The results show that 3B43 V7 product is able to capture the phase of the annual cycle of monthly mean precipitation, but the performance is not good for the amplitude, in particular over the Andes and Pacific regions owing to complex climatic and topographic conditions. In general, 3B43 V7 exhibits good performance in the low‐lying and plain Amazon, Orinoco and Caribbean regions. Over the Andes region, characterized by complex topography, overestimation errors are identified [root mean squared error (RMSE) ≥83.59 mm·month−1 and relative bias (BIAS) ≥4.69%], whereas the extremely wet rainfall regime of the Pacific region is largely underestimated (RMSE ≥253.52 mm ·month−1 and BIAS ≤−11.75%). These errors are greater during the wet seasons when the metrics reach worse scores than those reported in similar studies worldwide. Occurrence analyses showed that 3B43 V7 misses very frequent light rainfall events and less frequent but very heavy storms, which contribute to the overall underestimation (overestimation) observed over the Pacific (Andes) region. The error characteristics identified and quantified in this study confirm the well‐documented limitations of remote precipitation sensing and constitute a warning about major challenges that complex climatic and physiographic features can impose on satellite rainfall missions.

Columbia↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Assessment of Computational Fluid Dynamics (CFD) Models for Shock Boundary-Layer Interaction

A workshop on the computational fluid dynamics (CFD) prediction of shock boundary-layer interactions (SBLIs) was held at the 48th AIAA Aerospace Sciences Meeting. As part of the workshop numerous CFD analysts submitted solutions to four experimentally measured SBLIs. This paper describes the assessment of the CFD predictions. The assessment includes an uncertainty analysis of the experimental data, the definition of an error metric and the application of that metric to the CFD solutions. The CFD solutions provided very similar levels of error and in general it was difficult to discern clear trends in the data. For the Reynolds Averaged Navier-Stokes methods the choice of turbulence model appeared to be the largest factor in solution accuracy. Large-eddy simulation methods produced error levels similar to RANS methods but provided superior predictions of normal stresses.

DeBonis, James R.↗

Gleipnir: toward practical error analysis for Quantum programs

Practical error analysis is essential for the design, optimization, and evaluation of Noisy Intermediate-Scale Quantum(NISQ) computing. However, bounding errors in quantum programs is a grand challenge, because the effects of quantum errors depend on exponentially large quantum states. In this work, we present Gleipnir, a novel methodology toward practically computing verified error bounds in quantum programs. Gleipnir introduces the (ρ,δ)-diamond norm, an error metric constrained by a quantum predicate consisting of the approximate state ρ and its distance δ to the ideal state ρ. This predicate (ρ,δ) can be computed adaptively using tensor networks based on the Matrix Product States. Gleipnir features a lightweight logic for reasoning about error bounds in noisy quantum programs, based on the (ρ,δ)-diamond norm metric. Furthermore, our experimental results show that Gleipnir is able to efficiently generate tight error bounds for real-world quantum programs with 10 to 100 qubits, and can be used to evaluate the error mitigation performance of quantum compiler transformations.

Tao, Runzhou↗

Assessing Clouds Using Satellite Observations Through Three Generations of Global Atmosphere Models

Abstract Clouds are parameterized in climate models using quantities on the model grid‐scale to approximate the cloud cover and impact on radiation. Because of the complexity of processes involved with clouds, these parameterizations are one of the key challenges in climate modeling. Differences in parameterizations of clouds are among the main contributors to the spread in climate sensitivity across models. In this work, the clouds in three generations of an atmosphere model lineage are evaluated against satellite observations. Satellite simulators are used within the model to provide an appropriate comparison with individual satellite products. In some respects, especially the top‐of‐atmosphere cloud radiative effect, the models show generational improvements. The most recent generation, represented by two distinct branches of development, exhibits some regional regressions in the cloud representation; in particular the southern ocean shows a positive bias in cloud cover. The two branches of model development show how choices during model development, both structural and parametric, lead to different cloud climatologies. Several evaluation strategies are used to quantify the spatial errors in terms of the large‐scale circulation and the cloud structure. The Earth mover's distance is proposed as a useful error metric for the passive satellite data products that provide cloud‐top pressure‐optical depth histograms. The cloud errors identified here may contribute to the high climate sensitivity in the Community Earth System Model, version 2 and in the Energy Exascale Earth System Model, version 1.

54 ENVIRONMENTAL SCIENCES↗

A variational framework for residual-based adaptivity in neural PDE solvers and operator learning

Residual-based adaptive strategies are widely used in scientific machine learning yet remain largely heuristic. We introduce a variational framework that formalizes these methods through convex transformations of the residual, where different transformations correspond to distinct objective functionals. For instance, exponential weights target uniform error minimization, while linear weights recover quadratic error minimization. This perspective reveals adaptive weighting as a means of selecting sampling distributions that optimize a primal objective, directly linking discretization choices to error metrics. This principled approach yields three key benefits: it enables systematic design of adaptive schemes, reduces discretization error by lowering estimator variance, and enhances learning dynamics by improving gradient signal-to-noise ratio. Extending the framework to operator learning, we demonstrate substantial performance gains across diverse optimizers and architectures. Our results provide a theoretical perspective for residual-based adaptivity and establish a foundation for principled discretization and training.

97 MATHEMATICS AND COMPUTING↗

DCT quantization matrices visually optimized for individual images

This presentation describes how a vision model incorporating contrast sensitivity, contrast masking, and light adaptation is used to design visually optimal quantization matrices for Discrete Cosine Transform image compression. The Discrete Cosine Transform (DCT) underlies several image compression standards (JPEG, MPEG, H.261). The DCT is applied to 8x8 pixel blocks, and the resulting coefficients are quantized by division and rounding. The 8x8 'quantization matrix' of divisors determines the visual quality of the reconstructed image; the design of this matrix is left to the user. Since each DCT coefficient corresponds to a particular spatial frequency in a particular image region, each quantization error consists of a local increment or decrement in a particular frequency. After adjustments for contrast sensitivity, local light adaptation, and local contrast masking, this coefficient error can be converted to a just-noticeable-difference (jnd). The jnd's for different frequencies and image blocks can be pooled to yield a global perceptual error metric. With this metric, we can compute for each image the quantization matrix that minimizes the bit-rate for a given perceptual error, or perceptual error for a given bit-rate. Implementation of this system demonstrates its advantages over existing techniques. A unique feature of this scheme is that the quantization matrix is optimized for each individual image. This is compatible with the JPEG standard, which requires transmission of the quantization matrix.

Watson, Andrew B.↗

An Approach for the Assessment of System Upset Resilience

This report describes an approach for the assessment of upset resilience that is applicable to systems in general, including safety-critical, real-time systems. For this work, resilience is defined as the ability to preserve and restore service availability and integrity under stated conditions of configuration, functional inputs and environmental conditions. To enable a quantitative approach, we define novel system service degradation metrics and propose a new mathematical definition of resilience. These behavioral-level metrics are based on the fundamental service classification criteria of correctness, detectability, symmetry and persistence. This approach consists of a Monte-Carlo-based stimulus injection experiment, on a physical implementation or an error-propagation model of a system, to generate a system response set that can be characterized in terms of dimensional error metrics and integrated to form an overall measure of resilience. We expect this approach to be helpful in gaining insight into the error containment and repair capabilities of systems for a wide range of conditions.

Torres-Pomales, Wilfredo↗

Calibrating Microscopic Car-Following Models for Adaptive Cruise Control Vehicles: Multiobjective Approach

Adaptive cruise control (ACC) vehicles are the first step toward comprehensive vehicle automation. However, the impacts of such vehicles on the underlying traffic flow are not yet clear. Therefore, it is of interest to accurately model vehicle-level dynamics of commercially available ACC vehicles so that they may be used in further modeling efforts to quantify the impact of commercially available ACC vehicles on traffic flow. Importantly, not only model selection but also the calibration approach and error metric used for calibration are critical to accurately model ACC vehicle behavior. In this work, we explore the question of how to calibrate car following models to describe ACC vehicle dynamics. Specifically, we apply a multi-objective calibration approach to understand the tradeoff between calibrating model parameters to minimize speed error vs. spacing error. Three different car-following models are calibrated for data from six vehicles. The results are in line with recent literature and verify that targeting a low spacing error does not compromise the speed accuracy whether the opposite is not true for modeling ACC vehicle dynamics.

33 ADVANCED PROPULSION SYSTEMS↗

Worldwide benchmarking of cost-effective radiometers for direct and diffuse irradiance

Solar energy projects can benefit from direct normal irradiance (DNI) and diffuse horizontal irradiance (DHI) measurements during all project phases. Several commercial measurement systems for DNI and DHI are available. Sun trackers with pyranometers and pyrheliometers can provide highly accurate measurements but are often impractical in solar energy applications. For less expensive and more robust sensors, it is often unclear which accuracy can be expected under a project site's specific atmospheric conditions. We address this challenge through our dedicated experimental comparison of relevant sensor systems (rotating shadowband irradiometer [short RSI], Delta-T SPN1, EKO MS-90, PyranoCam, Sunto CaptPro, Kipp & Zonen CSD3) at up to six sites worldwide. The RSI systems (rRMSD 3 to 8.6%, DNI; 4.8 to 7.6%, DHI) and PyranoCam (rRMSD 2.6 to 5.2%, DNI; 4.4 to 5.8%, DHI) exhibit similar error metrics and are the most accurate systems in the test. Delta-T SPN1 and EKO MS-90 (rRMSD 6.8 to 15%, DNI; 10.6 to 20.1%, DHI) but especially Kipp & Zonen CSD3 and Sunto CaptPro show significant deviations (rRMSD 17.7 to 20%, DNI; 33 to 58%, DHI). We evaluate the influence of relevant atmospheric parameters on the sensors' accuracies by a rather unique measurement setup. MS-90's DNI errors depend on DNI itself, with overestimations for low reference DNI. The deviations of SPN1's DHI and DNI measurements increase sharply in situations with high circumsolar irradiance. Also CaptPro and CSD3's increased measurement errors are related to circumsolar irradiance. For RSI and PyranoCam, only moderate influences on the measurements are identified, indicating a general applicability of these instruments.

14 SOLAR ENERGY↗

Understanding Biases in Simulated Cloud Radiative Effects in E3SMv3

This study systematically investigates biases in cloud radiative effects (CREs) within the recently released Energy Exascale Earth System Model version 3 (E3SMv3). Compared to its previous version (E3SMv2), E3SMv3 shows excessively strong shortwave CRE over tropical and subtropical oceans, the Southern Ocean, and the Northern Hemisphere storm tracks, which is primarily caused by an overestimation of optically intermediate low clouds. The model also displays excessive longwave CRE over the Maritime Continent and other tropical deep convection regions, resulting from an overestimation of optically thick high clouds. Implementing the Predicted Particle Properties (P3) scheme for stratiform clouds played the most significant role in these cloud changes. In addition, using a double-moment scheme for convective clouds contributed to the increase in intermediate low clouds and optically thick clouds in tropical deep convection regions. Error metrics for total cloud amount (E TCA ), cloud properties (E ctp-τ ), and cloud properties weighted by their SW and LW radiative impacts (E SW and E LW ) indicate that E3SMv3's ability to reproduce observed cloud radiative effect of low, middle, and high clouds has not been improved compared to E3SMv2. Nevertheless, the performance of E3SMv3 remains well within the spread of CMIP6 models, with E SW and E LW values smaller than those in most CMIP6 models. This study underscores the importance of integrating diverse satellite observations for robust cloud evaluation and using cloud-radiation relationships as consistency check for model errors. It suggests that further model developments focus on improving cloud microphysics and their interactions with radiation.

Environmental sciences↗

Active learning for SNAP interatomic potentials via Bayesian predictive uncertainty

Bayesian inference with a simple Gaussian error model is used to efficiently compute prediction variances for energies, forces, and stresses in the linear SNAP interatomic potential. Here, the prediction variance is shown to have a strong correlation with the absolute error over approximately 24 orders of magnitude. Using this prediction variance, an active learning algorithm is constructed to iteratively train a potential by selecting the structures with the most uncertain properties from a pool of candidate structures. The relative importance of the energy, force, and stress errors in the objective function is shown to have a strong impact upon the trajectory of their respective net error metrics when running the active learning algorithm. Batched training of different batch sizes is also tested against singular structure updates, and it is found that batches can be used to significantly reduce the number of retraining steps required with only minor impact on the active learning trajectory.

97 MATHEMATICS AND COMPUTING↗

On the effectiveness of neural operators at zero-shot weather downscaling

Machine-learning (ML) methods have shown great potential for weather downscaling. These data-driven approaches provide a more efficient alternative for producing high-resolution weather datasets and forecasts compared to physics-based numerical simulations. Neural operators, which learn solution operators for a family of partial differential equations, have shown great success in scientific ML applications involving physics-driven datasets. Neural operators are grid-resolution-invariant and are often evaluated on higher grid resolutions than they are trained on, i.e., zero-shot super-resolution. Given their promising zero-shot super-resolution performance on dynamical systems emulation, we present a critical investigation of their zero-shot weather downscaling capabilities, which is when models are tasked with producing high-resolution outputs using higher upsampling factors than are seen during training. To this end, we create two realistic downscaling experiments with challenging upsampling factors (e.g., 8x and 15x) across data from different simulations: the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) and the Wind Integration National Dataset Toolkit. While neural operator-based downscaling models perform better than interpolation and a simple convolutional baseline, we show the surprising performance of an approach that combines a powerful transformer-based model with parameter-free interpolation at zero-shot weather downscaling. We find that this Swin-Transformer-based approach mostly outperforms models with neural operator layers in terms of average error metrics, whereas an Enhanced Super-Resolution Generative Adversarial Network-based approach is better than most models in terms of capturing the physics of the ground truth data. We suggest their use in future work as strong baselines.

17 WIND ENERGY↗

Machine learning surrogates for ion energy–angle distributions in thermal and RF plasma sheaths

Ion energy–angle distributions (IEADs) at material surfaces are a critical input for plasma–material interaction (PMI) studies in fusion devices, yet they are computationally expensive to obtain using particle-in-cell (PIC) simulations. In this work, we develop a machine learning surrogate based on a deep deconvolutional neural network (DDeCNN) trained on large databases generated with the hPIC2 code. The surrogate is capable of reconstructing IEADs from sheath parameters for both thermal and radio-frequency (RF) plasmas, including cases with multiple ion species. Across thousands of test cases, the model achieves high accuracy, with over 97 % of predictions classified as good or average based on standard error metrics (MAE, MSE, L2). Even in the more challenging RF and multi-species regimes, the surrogate reliably captures the multi-peak structure of PIC results. Once trained, the surrogate produces IEADs in milliseconds on a common workstation, yielding speedups of six to seven orders of magnitude compared with running a full PIC simulation. This computational gain enables dense parameter scans and direct coupling of IEAD predictions with PMI and erosion models on whole-device scales in fusion-relevant conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Large Ensemble Diagnostic Evaluation of Hydrologic Parameter Uncertainty in the Community Land Model Version 5 (CLM5)

Abstract Land surface models such as the Community Land Model version 5 (CLM5) seek to enhance understanding of terrestrial hydrology and aid in the evaluation of anthropogenic and climate change impacts. However, the effects of parametric uncertainty on CLM5 hydrologic predictions across regions, timescales, and flow regimes have yet to be explored in detail. The common use of the default hydrologic model parameters in CLM5 risks generating streamflow predictions that may lead to incorrect inferences for important dynamics and/or extremes. In this study, we benchmark CLM5 streamflow predictions relative to the commonly employed default hydrologic parameters for 464 headwater basins over the conterminous United States (CONUS). We evaluate baseline CLM5 default parameter performance relative to a large (1,307) Latin Hypercube Sampling‐based diagnostic comparison of streamflow prediction skill using over 20 error measures. We provide a global sensitivity analysis that clarifies the significant spatial variations in parametric controls for CLM5 streamflow predictions across regions, temporal scales, and error metrics of interest. The baseline CLM5 shows relatively moderate to poor streamflow prediction skill in several CONUS regions, especially the arid Southwest and Central U.S. Hydrologic parameter uncertainty strongly affects CLM5 streamflow predictions, but its impacts vary in complex ways across U.S. regions, timescales, and flow regimes. Overall, CLM5's surface runoff and soil water parameters have the largest effects on simulated high flows, while canopy water and evaporation parameters have the most significant effects on the water balance.

54 ENVIRONMENTAL SCIENCES↗