Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error metric”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Improved Representation of Horizontal Variability and Turbulence in Mesoscale Simulations of an Extended Cold-Air Pool Event

Abstract Cold-air pools (CAPs), or stable atmospheric boundary layers that form within topographic basins, are associated with poor air quality, hazardous weather, and low wind energy output. Accurate prediction of CAP dynamics presents a challenge for mesoscale forecast models in part because CAPs occur in regions of complex terrain, where traditional turbulence parameterizations may not be appropriate. This study examines the effects of the planetary boundary layer (PBL) scheme and horizontal diffusion treatment on CAP prediction in the Weather Research and Forecasting (WRF) Model. Model runs with a one-dimensional (1D) PBL scheme and Smagorinsky-like horizontal diffusion are compared with runs that use a new three-dimensional (3D) PBL scheme to calculate turbulent fluxes. Simulations are completed in a nested configuration with 3-km/750-m horizontal grid spacing over a 10-day case study in the Columbia River basin, and results are compared with observations from the Second Wind Forecast Improvement Project. Using event-averaged error metrics, potential temperature and wind speed errors are shown to decrease both with increased horizontal grid resolution and with improved treatment of horizontal diffusion over steep terrain. The 3D PBL scheme further reduces errors relative to a standard 1D PBL approach. Error reduction is accentuated during CAP erosion, when turbulent mixing plays a more dominant role in the dynamics. Last, the 3D PBL scheme is shown to reduce near-surface overestimates of turbulence kinetic energy during the CAP event. The sensitivity of turbulence predictions to the master length-scale formulation in the 3D PBL parameterization is also explored. Significance Statement In this article, we demonstrate how a new framework for modeling atmospheric turbulence improves cold pool predictions, using a case study from January 2017 in the Columbia River basin (U.S. Pacific Northwest). Cold pools are regions of cold, stagnant air that form within valleys or basins, and improved forecasts could help to mitigate the risks they pose to air quality, transportation, and wind energy production. For the chosen case study, our tests show a reduction in temperature and wind speed errors by up to a factor of 2–3 relative to standard model options. These results strongly motivate continued development of the framework as well as its application to other complex weather events.

17 WIND ENERGY↗

A North Sea in Situ Evaluation of the Fitch Wind Farm Parameterization Within the Mellor-Yamada-Nakanishi-Niino and 3D Planetary Boundary Layer Schemes

Wind resource assessments and wind power forecasts that account for wind farm wakes are sensitive to the choice of planetary boundary layer (PBL) scheme. This work compares the one-dimensional Mellor-Yamada-Nakanishi-Niino (MYNN) PBL scheme with a three-dimensional PBL (3DPBL) scheme, evaluating predictions made with both schemes against two sets of North Sea in situ observations of wind farm wakes. The optimal PBL scheme varies based on the observations (FINO1 tower vs. aircraft), the quantity of interest (wind speed vs. turbulence kinetic energy [TKE]), and the error metric (bias, centered root mean square error [cRMSE], R2, and earth mover's distance [EMD]). Whereas 3DPBL wind speeds outperform MYNN wind speeds with respect to the cRMSE at the FINO1 site located at a single point within the turbine rotor layer, 3DPBL TKE bias is larger than MYNN TKE bias when compared to aircraft observations taken 100 m above a wind farm. Wind speeds in the aircraft region are ambiguous with regard to which PBL scheme is optimal. Aircraft MYNN wind speeds outperform 3DPBL wind speeds with respect to R2 and cRMSE but underperform with respect to bias and EMD. Future evaluations across broader temporal and spatial scales may offer further insight into model differences.

17 WIND ENERGY↗

Model Calibration with Markov Chain Monte Carlo Tutorial

The purpose of this tutorial is to demonstrate how to use Markov chain Monte Carlo (MCMC) to calibrate a model. By calibration, we mean the selection of model parameters (and, when relevant, structures). A common goal in model development and diagnostics is calibration, or the identification of model structures and parameters which are consistent with data. While models can be calibrated through hand-tuning parameters or minimizing simple error metrics such as root-mean-square-error (RMSE), these approaches can underrepresent the probabilistic nature of the data-generating process, as well as the potential for multiple model configurations to be consistent with the data. Probabilistic uncertainty quantification, which is the topic of this notebook, can address these concerns. This tutorial is presented as an appendix to the e-book: Addressing Uncertainty in MultiSector Dynamics Research.

Markov chain Monte Carlo↗

Evaluation of Two Crew Module Boilerplate Tests Using Newly Developed Calibration Metrics

The paper discusses a application of multi-dimensional calibration metrics to evaluate pressure data from water drop tests of the Max Launch Abort System (MLAS) crew module boilerplate. Specifically, three metrics are discussed: 1) a metric to assess the probability of enveloping the measured data with the model, 2) a multi-dimensional orthogonality metric to assess model adequacy between test and analysis, and 3) a prediction error metric to conduct sensor placement to minimize pressure prediction errors. Data from similar (nearly repeated) capsule drop tests shows significant variability in the measured pressure responses. When compared to expected variability using model predictions, it is demonstrated that the measured variability cannot be explained by the model under the current uncertainty assumptions.

Horta, Lucas G.↗

Multiresolution Quantum Chemistry: Nonlinear Response Properties at the Basis Set Limit

We benchmark the accuracy of Dunning correlation-consistent Gaussian basis sets for computing frequencydependent second-order hyperpolarizabilities relevant to second-harmonic generation (SHG), using multiresolution analysis (MRA) as a reference. Basis set errors are analyzed using a unit-sphere representation of the effective hyperpolarizability vector, enabling direct assessment of directional error structure. We introduce a relative RMS total error metric that integrates directional deviations over the unit sphere and complement it with signed projection errors that distinguish over- and underestimation. Unsupervised clustering based on these signed directional metrics reveals four distinct convergence behaviors across a set of 68 molecules. Unitsphere visualizations of representative systems show that basis set errors are often highly anisotropic and localized along specific bond directions, even when global error measures appear small. Doubly augmented basis sets consistently outperform singly augmented ones, and core-polarization functions are required for uniform convergence in second-row systems. Overall, this work demonstrates that directional analysis combined with clustering provides a robust framework for understanding basis set convergence in nonlinear optical response properties.

Basis sets↗

Ground Validation of TRMM 3B43 V7 Precipitation Estimates Over Colombia. Part I: Monthly and Seasonal Timescales

In this study, we validate precipitation estimates remotely sensed by the Tropical Rainfall Measuring Mission (TRMM) at monthly and seasonal timescales, during the period 1998–2015, by calculating and analyzing diverse error metrics between the 3B43 V7 product and in situ measurements from 1,180 rain gauges over Colombia, of which at least 987 are fully independent of TRMM. We explore the existence of spatiotemporal patterns to assess the performance of 3B43 V7 over the five major natural regions of Colombia: Caribbean, Pacific, Andes, Orinoco and Amazon. The results show that 3B43 V7 product is able to capture the phase of the annual cycle of monthly mean precipitation, but the performance is not good for the amplitude, in particular over the Andes and Pacific regions owing to complex climatic and topographic conditions. In general, 3B43 V7 exhibits good performance in the low‐lying and plain Amazon, Orinoco and Caribbean regions. Over the Andes region, characterized by complex topography, overestimation errors are identified [root mean squared error (RMSE) ≥83.59 mm·month−1 and relative bias (BIAS) ≥4.69%], whereas the extremely wet rainfall regime of the Pacific region is largely underestimated (RMSE ≥253.52 mm ·month−1 and BIAS ≤−11.75%). These errors are greater during the wet seasons when the metrics reach worse scores than those reported in similar studies worldwide. Occurrence analyses showed that 3B43 V7 misses very frequent light rainfall events and less frequent but very heavy storms, which contribute to the overall underestimation (overestimation) observed over the Pacific (Andes) region. The error characteristics identified and quantified in this study confirm the well‐documented limitations of remote precipitation sensing and constitute a warning about major challenges that complex climatic and physiographic features can impose on satellite rainfall missions.

Columbia↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Assessment of Computational Fluid Dynamics (CFD) Models for Shock Boundary-Layer Interaction

A workshop on the computational fluid dynamics (CFD) prediction of shock boundary-layer interactions (SBLIs) was held at the 48th AIAA Aerospace Sciences Meeting. As part of the workshop numerous CFD analysts submitted solutions to four experimentally measured SBLIs. This paper describes the assessment of the CFD predictions. The assessment includes an uncertainty analysis of the experimental data, the definition of an error metric and the application of that metric to the CFD solutions. The CFD solutions provided very similar levels of error and in general it was difficult to discern clear trends in the data. For the Reynolds Averaged Navier-Stokes methods the choice of turbulence model appeared to be the largest factor in solution accuracy. Large-eddy simulation methods produced error levels similar to RANS methods but provided superior predictions of normal stresses.

DeBonis, James R.↗

Gleipnir: toward practical error analysis for Quantum programs

Practical error analysis is essential for the design, optimization, and evaluation of Noisy Intermediate-Scale Quantum(NISQ) computing. However, bounding errors in quantum programs is a grand challenge, because the effects of quantum errors depend on exponentially large quantum states. In this work, we present Gleipnir, a novel methodology toward practically computing verified error bounds in quantum programs. Gleipnir introduces the (ρ,δ)-diamond norm, an error metric constrained by a quantum predicate consisting of the approximate state ρ and its distance δ to the ideal state ρ. This predicate (ρ,δ) can be computed adaptively using tensor networks based on the Matrix Product States. Gleipnir features a lightweight logic for reasoning about error bounds in noisy quantum programs, based on the (ρ,δ)-diamond norm metric. Furthermore, our experimental results show that Gleipnir is able to efficiently generate tight error bounds for real-world quantum programs with 10 to 100 qubits, and can be used to evaluate the error mitigation performance of quantum compiler transformations.

Tao, Runzhou↗

Assessing Clouds Using Satellite Observations Through Three Generations of Global Atmosphere Models

Abstract Clouds are parameterized in climate models using quantities on the model grid‐scale to approximate the cloud cover and impact on radiation. Because of the complexity of processes involved with clouds, these parameterizations are one of the key challenges in climate modeling. Differences in parameterizations of clouds are among the main contributors to the spread in climate sensitivity across models. In this work, the clouds in three generations of an atmosphere model lineage are evaluated against satellite observations. Satellite simulators are used within the model to provide an appropriate comparison with individual satellite products. In some respects, especially the top‐of‐atmosphere cloud radiative effect, the models show generational improvements. The most recent generation, represented by two distinct branches of development, exhibits some regional regressions in the cloud representation; in particular the southern ocean shows a positive bias in cloud cover. The two branches of model development show how choices during model development, both structural and parametric, lead to different cloud climatologies. Several evaluation strategies are used to quantify the spatial errors in terms of the large‐scale circulation and the cloud structure. The Earth mover's distance is proposed as a useful error metric for the passive satellite data products that provide cloud‐top pressure‐optical depth histograms. The cloud errors identified here may contribute to the high climate sensitivity in the Community Earth System Model, version 2 and in the Energy Exascale Earth System Model, version 1.

54 ENVIRONMENTAL SCIENCES↗

A variational framework for residual-based adaptivity in neural PDE solvers and operator learning

Residual-based adaptive strategies are widely used in scientific machine learning yet remain largely heuristic. We introduce a variational framework that formalizes these methods through convex transformations of the residual, where different transformations correspond to distinct objective functionals. For instance, exponential weights target uniform error minimization, while linear weights recover quadratic error minimization. This perspective reveals adaptive weighting as a means of selecting sampling distributions that optimize a primal objective, directly linking discretization choices to error metrics. This principled approach yields three key benefits: it enables systematic design of adaptive schemes, reduces discretization error by lowering estimator variance, and enhances learning dynamics by improving gradient signal-to-noise ratio. Extending the framework to operator learning, we demonstrate substantial performance gains across diverse optimizers and architectures. Our results provide a theoretical perspective for residual-based adaptivity and establish a foundation for principled discretization and training.

97 MATHEMATICS AND COMPUTING↗

DCT quantization matrices visually optimized for individual images

This presentation describes how a vision model incorporating contrast sensitivity, contrast masking, and light adaptation is used to design visually optimal quantization matrices for Discrete Cosine Transform image compression. The Discrete Cosine Transform (DCT) underlies several image compression standards (JPEG, MPEG, H.261). The DCT is applied to 8x8 pixel blocks, and the resulting coefficients are quantized by division and rounding. The 8x8 'quantization matrix' of divisors determines the visual quality of the reconstructed image; the design of this matrix is left to the user. Since each DCT coefficient corresponds to a particular spatial frequency in a particular image region, each quantization error consists of a local increment or decrement in a particular frequency. After adjustments for contrast sensitivity, local light adaptation, and local contrast masking, this coefficient error can be converted to a just-noticeable-difference (jnd). The jnd's for different frequencies and image blocks can be pooled to yield a global perceptual error metric. With this metric, we can compute for each image the quantization matrix that minimizes the bit-rate for a given perceptual error, or perceptual error for a given bit-rate. Implementation of this system demonstrates its advantages over existing techniques. A unique feature of this scheme is that the quantization matrix is optimized for each individual image. This is compatible with the JPEG standard, which requires transmission of the quantization matrix.

Watson, Andrew B.↗

An Approach for the Assessment of System Upset Resilience

This report describes an approach for the assessment of upset resilience that is applicable to systems in general, including safety-critical, real-time systems. For this work, resilience is defined as the ability to preserve and restore service availability and integrity under stated conditions of configuration, functional inputs and environmental conditions. To enable a quantitative approach, we define novel system service degradation metrics and propose a new mathematical definition of resilience. These behavioral-level metrics are based on the fundamental service classification criteria of correctness, detectability, symmetry and persistence. This approach consists of a Monte-Carlo-based stimulus injection experiment, on a physical implementation or an error-propagation model of a system, to generate a system response set that can be characterized in terms of dimensional error metrics and integrated to form an overall measure of resilience. We expect this approach to be helpful in gaining insight into the error containment and repair capabilities of systems for a wide range of conditions.

Torres-Pomales, Wilfredo↗

Calibrating Microscopic Car-Following Models for Adaptive Cruise Control Vehicles: Multiobjective Approach

Adaptive cruise control (ACC) vehicles are the first step toward comprehensive vehicle automation. However, the impacts of such vehicles on the underlying traffic flow are not yet clear. Therefore, it is of interest to accurately model vehicle-level dynamics of commercially available ACC vehicles so that they may be used in further modeling efforts to quantify the impact of commercially available ACC vehicles on traffic flow. Importantly, not only model selection but also the calibration approach and error metric used for calibration are critical to accurately model ACC vehicle behavior. In this work, we explore the question of how to calibrate car following models to describe ACC vehicle dynamics. Specifically, we apply a multi-objective calibration approach to understand the tradeoff between calibrating model parameters to minimize speed error vs. spacing error. Three different car-following models are calibrated for data from six vehicles. The results are in line with recent literature and verify that targeting a low spacing error does not compromise the speed accuracy whether the opposite is not true for modeling ACC vehicle dynamics.

33 ADVANCED PROPULSION SYSTEMS↗

Worldwide benchmarking of cost-effective radiometers for direct and diffuse irradiance

Solar energy projects can benefit from direct normal irradiance (DNI) and diffuse horizontal irradiance (DHI) measurements during all project phases. Several commercial measurement systems for DNI and DHI are available. Sun trackers with pyranometers and pyrheliometers can provide highly accurate measurements but are often impractical in solar energy applications. For less expensive and more robust sensors, it is often unclear which accuracy can be expected under a project site's specific atmospheric conditions. We address this challenge through our dedicated experimental comparison of relevant sensor systems (rotating shadowband irradiometer [short RSI], Delta-T SPN1, EKO MS-90, PyranoCam, Sunto CaptPro, Kipp & Zonen CSD3) at up to six sites worldwide. The RSI systems (rRMSD 3 to 8.6%, DNI; 4.8 to 7.6%, DHI) and PyranoCam (rRMSD 2.6 to 5.2%, DNI; 4.4 to 5.8%, DHI) exhibit similar error metrics and are the most accurate systems in the test. Delta-T SPN1 and EKO MS-90 (rRMSD 6.8 to 15%, DNI; 10.6 to 20.1%, DHI) but especially Kipp & Zonen CSD3 and Sunto CaptPro show significant deviations (rRMSD 17.7 to 20%, DNI; 33 to 58%, DHI). We evaluate the influence of relevant atmospheric parameters on the sensors' accuracies by a rather unique measurement setup. MS-90's DNI errors depend on DNI itself, with overestimations for low reference DNI. The deviations of SPN1's DHI and DNI measurements increase sharply in situations with high circumsolar irradiance. Also CaptPro and CSD3's increased measurement errors are related to circumsolar irradiance. For RSI and PyranoCam, only moderate influences on the measurements are identified, indicating a general applicability of these instruments.

14 SOLAR ENERGY↗

Understanding Biases in Simulated Cloud Radiative Effects in E3SMv3

This study systematically investigates biases in cloud radiative effects (CREs) within the recently released Energy Exascale Earth System Model version 3 (E3SMv3). Compared to its previous version (E3SMv2), E3SMv3 shows excessively strong shortwave CRE over tropical and subtropical oceans, the Southern Ocean, and the Northern Hemisphere storm tracks, which is primarily caused by an overestimation of optically intermediate low clouds. The model also displays excessive longwave CRE over the Maritime Continent and other tropical deep convection regions, resulting from an overestimation of optically thick high clouds. Implementing the Predicted Particle Properties (P3) scheme for stratiform clouds played the most significant role in these cloud changes. In addition, using a double-moment scheme for convective clouds contributed to the increase in intermediate low clouds and optically thick clouds in tropical deep convection regions. Error metrics for total cloud amount (E TCA ), cloud properties (E ctp-τ ), and cloud properties weighted by their SW and LW radiative impacts (E SW and E LW ) indicate that E3SMv3's ability to reproduce observed cloud radiative effect of low, middle, and high clouds has not been improved compared to E3SMv2. Nevertheless, the performance of E3SMv3 remains well within the spread of CMIP6 models, with E SW and E LW values smaller than those in most CMIP6 models. This study underscores the importance of integrating diverse satellite observations for robust cloud evaluation and using cloud-radiation relationships as consistency check for model errors. It suggests that further model developments focus on improving cloud microphysics and their interactions with radiation.

Environmental sciences↗

Active learning for SNAP interatomic potentials via Bayesian predictive uncertainty

Bayesian inference with a simple Gaussian error model is used to efficiently compute prediction variances for energies, forces, and stresses in the linear SNAP interatomic potential. Here, the prediction variance is shown to have a strong correlation with the absolute error over approximately 24 orders of magnitude. Using this prediction variance, an active learning algorithm is constructed to iteratively train a potential by selecting the structures with the most uncertain properties from a pool of candidate structures. The relative importance of the energy, force, and stress errors in the objective function is shown to have a strong impact upon the trajectory of their respective net error metrics when running the active learning algorithm. Batched training of different batch sizes is also tested against singular structure updates, and it is found that batches can be used to significantly reduce the number of retraining steps required with only minor impact on the active learning trajectory.

97 MATHEMATICS AND COMPUTING↗

On the effectiveness of neural operators at zero-shot weather downscaling

Machine-learning (ML) methods have shown great potential for weather downscaling. These data-driven approaches provide a more efficient alternative for producing high-resolution weather datasets and forecasts compared to physics-based numerical simulations. Neural operators, which learn solution operators for a family of partial differential equations, have shown great success in scientific ML applications involving physics-driven datasets. Neural operators are grid-resolution-invariant and are often evaluated on higher grid resolutions than they are trained on, i.e., zero-shot super-resolution. Given their promising zero-shot super-resolution performance on dynamical systems emulation, we present a critical investigation of their zero-shot weather downscaling capabilities, which is when models are tasked with producing high-resolution outputs using higher upsampling factors than are seen during training. To this end, we create two realistic downscaling experiments with challenging upsampling factors (e.g., 8x and 15x) across data from different simulations: the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) and the Wind Integration National Dataset Toolkit. While neural operator-based downscaling models perform better than interpolation and a simple convolutional baseline, we show the surprising performance of an approach that combines a powerful transformer-based model with parameter-free interpolation at zero-shot weather downscaling. We find that this Swin-Transformer-based approach mostly outperforms models with neural operator layers in terms of average error metrics, whereas an Enhanced Super-Resolution Generative Adversarial Network-based approach is better than most models in terms of capturing the physics of the ground truth data. We suggest their use in future work as strong baselines.

17 WIND ENERGY↗