Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Highly accurate and constrained density functional obtained with differentiable programming

Using an end-to-end differentiable implementation of the Kohn-Sham self-consistent field equations, we obtain a highly accurate neural network–based exchange and correlation (XC) functional of the electronic density. The functional is optimized using information on both energy and density while exact constraints are enforced through an appropriate neural network architecture. Here we evaluate our model against different families of XC approximations and show that at the meta-GGA level our functional exhibits unprecedented accuracy for both energy and density predictions. For nonempirical functionals, there is a strong linear correlation between energy and density errors. We use this correlation to define an XC functional quality metric that includes both energy and density errors, leading to an improved way to rank different approximations.

36 MATERIALS SCIENCE↗

Similarity Metrics for Closed Loop Dynamic Systems

To what extent and in what ways can two closed-loop dynamic systems be said to be "similar?" This question arises in a wide range of dynamic systems modeling and control system design applications. For example, bounds on error models are fundamental to the controller optimization with modern control design methods. Metrics such as the structured singular value are direct measures of the degree to which properties such as stability or performance are maintained in the presence of specified uncertainties or variations in the plant model. Similarly, controls-related areas such as system identification, model reduction, and experimental model validation employ measures of similarity between multiple realizations of a dynamic system. Each area has its tools and approaches, with each tool more or less suited for one application or the other. Similarity in the context of closed-loop model validation via flight test is subtly different from error measures in the typical controls oriented application. Whereas similarity in a robust control context relates to plant variation and the attendant affect on stability and performance, in this context similarity metrics are sought that assess the relevance of a dynamic system test for the purpose of validating the stability and performance of a "similar" dynamic system. Similarity in the context of system identification is much more relevant than are robust control analogies in that errors between one dynamic system (the test article) and another (the nominal "design" model) are sought for the purpose of bounding the validity of a model for control design and analysis. Yet system identification typically involves open-loop plant models which are independent of the control system (with the exception of limited developments in closed-loop system identification which is nonetheless focused on obtaining open-loop plant models from closed-loop data). Moreover the objectives of system identification are not the same as a flight test and hence system identification error metrics are not directly relevant. In applications such as launch vehicles where the open loop plant is unstable it is similarity of the closed-loop system dynamics of a flight test that are relevant.

Whorton, Mark S.↗

A new neural net approach to robot 3D perception and visuo-motor coordination

A novel neural network approach to robot hand-eye coordination is presented. The approach provides a true sense of visual error servoing, redundant arm configuration control for collision avoidance, and invariant visuo-motor learning under gazing control. A 3-D perception network is introduced to represent the robot internal 3-D metric space in which visual error servoing and arm configuration control are performed. The arm kinematic network performs the bidirectional association between 3-D space arm configurations and joint angles, and enforces the legitimate arm configurations. The arm kinematic net is structured by a radial-based competitive and cooperative network with hierarchical self-organizing learning. The main goal of the present work is to demonstrate that the neural net representation of the robot 3-D perception net serves as an important intermediate functional block connecting robot eyes and arms.

Lee, Sukhan↗

An experimental evaluation of error seeding as a program validation technique

A previously reported experiment in error seeding as a program validation technique is summarized. The experiment was designed to test the validity of three assumptions on which the alleged effectiveness of error seeding is based. Errors were seeded into 17 functionally identical but independently programmed Pascal programs in such a way as to produce 408 programs, each with one seeded error. Using mean time to failure as a metric, results indicated that it is possible to generate seeded errors that are arbitrarily but not equally difficult to locate. Examination of indigenous errors demonstrated that these are also arbitrarily difficult to locate. These two results support the assumption that seeded and indigenous errors are approximately equally difficult to locate. However, the assumption that, for each type of error, all errors are equally difficult to locate was not borne out. Finally, since a seeded error occasionally corrected an indigenous error, the assumption that errors do not interfere with each other was proven wrong. Error seeding can be made useful by taking these results into account in modifying the underlying model.

Knight, J. C.↗

Towards Scaling Law Analysis For Spatiotemporal Weather Data

Compute-optimal scaling laws are relatively well studied for NLP and CV, where objectives are typically single-step and targets are comparatively homogeneous. Weather forecasting is harder to characterize in the same framework: autoregressive rollouts compound errors over long horizons, outputs couple many physical channels with disparate scales and predictability, and globally pooled test metrics can disagree sharply with per-channel, late-lead behavior implied by short-horizon training. We extend neural scaling analysis for autoregressive weather forecasting from single-step training loss to long rollouts and per-channel metrics. We quantify (1) how prediction error is distributed across channels and how its growth rate evolves with forecast horizon, (2) if power law scaling holds for test error, relative to rollout length when error is pooled globally, and (3) how that fit varies jointly with horizon and channel for parameter, data, and compute-based scaling axes. We find strong cross-channel and cross-horizon heterogeneity: pooled scaling can look favorable while many channels degrade at late leads. We discuss implications for weighted objectives, horizon-aware curricula, and resource allocation across outputs.

Kiefer Jr, Alexander [ORNL] (ORCID:000000025398874↗

Probabilistic Day-Ahead Forecasting Using an Analog Ensemble Approach for Wind Farm Grid Services

Wind resource assessment and wind power forecasting are used in research and industry to anticipate future power output at scales ranging from individual wind turbines to entire wind farms. Probabilistic day-ahead wind forecasting is useful for anticipating how a wind farm could potentially participate in the day-ahead market by providing upper and lower bounds for expected power generation, thus informing grid operators of its uncertainty. Understanding this uncertainty is part of a larger project focused on building a platform that combines efforts in weather forecasting, aerodynamic and economic modeling to create maximum value of a wind plant to better provide services to the grid. This effort is also known as the Atmosphere to Electrons to Grid (A2E2G) project. One method for producing a probabilistic forecast is through the analog ensemble approach (Delle Monache et al., 2011). This method leverages historical forecasts and their corresponding observations as a training data set from which future forecasts can be made. For some future forecast, the most similar historical forecasts (analogs) are identified on a regular time basis such as once per a 3-hour window. The most similar analogs, based on a metric such as root mean square error (RMSE), are recorded and their corresponding verifying observations are used as an ensemble member for this future forecast. Prior work in this area demonstrates improvements over raw Numerical Weather Prediction (NWP) forecasts and shows skill similar to techniques such as logistic regression and machine learning (Delle Monache et al., 2013; Alessandrini et al., 2015). Here, we take the High-Resolution Rapid Refresh model (HRRR) day-ahead forecast (0-36 hours) to create a probabilistic day-ahead forecast using an analog ensemble approach. The HRRR has an hourly temporal resolution, with a spatial resolution of 3 km. The 12 UTC HRRR model run is downloaded every day for one year from August 2019 - July 2020, with the first 11 months serving as a bank of analogs from which the forecasting algorithm can create a probabilistic forecast. Once downloaded, the original HRRR forecast is temporally interpolated to 5-minutes, aligning with both the temporal resolution of the observations as well as the timescale relevant for day-ahead power forecasts. The forecast is validated at the M2 tower at the Flatirons Campus of the National Renewable Energy Laboratory (NREL) at a typical wind turbine height of 80 m. Variables such as wind speed, wind direction, and turbulence intensity are incorporated into the probabilistic forecast model and weighted according to their relative importance to the forecast. Based on metrics such as mean bias error (MBE), mean absolute error (MAE), and root mean square error, the analog ensemble forecast outperforms the raw HRRR forecast during the testing period of July 2020. Figure 1 illustrates an example day-ahead forecast compared against the verifying observations. The general variability and ramps are captured throughout the day, with potential to further improve the analog ensemble model through machine learning techniques.

numerical weather prediction↗

Uncertainty Quantification for Empirical X-59 Sonic Boom Loudness Levels

Estimates of the total uncertainty for empirically determined loudness levels are documented when GRS (Ground Recording System) noise monitors are used to record X 59 sonic boom waveforms. The total uncertainty is characterized by combining nine different sources of uncertainty that may affect the apparent gain of the measurement chain. These uncertainty estimates are presented as expected measurement error relative to the true loudness level, and separate error estimates are provided for eight different noise metrics in which NASA has interest. The behavior of the Perceived Level (PL) metric is studied within the report body, while the total uncertainties for the seven other noise metrics are summarized in appendices for brevity. The effects of four sources of uncertainty are estimated simply from information found on hardware specification sheets provided by the manufacturer. However, mock acoustic recordings are created to estimate the effects of other sources because those effects are expected to induce spectral coloration, so they may vary with noise metric type and sound level. These sources are not well modeled by simple gain adjustments. Importantly, measurement error is computable when processing mock recordings since the true levels are knowable, which is not the case when processing data recorded in the field. Specifically, the true levels are knowable because the components of the mock recordings are separable – e.g., loudness levels of booms can be computed with or without superimposed background noise. Mock acoustic recordings also have the benefit of allowing analysis of sonic booms from vehicles that are not yet flying, like the X-59, since the mock recordings are created by combining vehicle-specific predicted waveforms with other audio sources. The estimates of total measurement error are documented as a function of the signal-to-noise ratio (SNR) of the loudness level, where the corrected SNR is computed while accounting for the effects of the method that is used to correct for background noise contamination when computing the noise metric values. The corrected SNR calculations used here can be applied to both mock recordings and in-field measurements, so the uncertainty of in-field recordings can be found using pre-computed lookup tables that identify the relationship between metric type, corrected SNR, and the expected measurement error.

Sonic Boom↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Quantifying the effects of mixing state on aerosol optical properties

Abstract. Calculations of the aerosol direct effect on climate rely on simulated aerosol fields. The model representation of aerosol mixing state potentially introduces large uncertainties into these calculations, since the simulated aerosol optical properties are sensitive to mixing state. In this study, we systematically quantified the impact of aerosol mixing state on aerosol optical properties using an ensemble of 1800 aerosol populations from particle-resolved simulations as a basis for Mie calculations for optical properties. Assuming the aerosol to be internally mixed within prescribed size bins caused overestimations of aerosol absorptivity and underestimations of aerosol scattering. Together, these led to errors in the populations' single scattering albedo of up to −22.3 % with a median of −0.9 %. The mixing state metric χ proved useful in relating errors in the volume absorption coefficient, the volume scattering coefficient and the single scattering albedo to the degree of internally mixing of the aerosol, with larger errors being associated with more external mixtures. At the same time, a range of errors existed for any given value of χ. We attributed this range to the extent to which the internal mixture assumption distorted the particles' black carbon content and the refractive index of the particle coatings. Both can vary for populations with the same value of χ. These results are further evidence of the important yet complicated role of mixing state in calculating aerosol optical properties.

54 ENVIRONMENTAL SCIENCES↗

High-accuracy Mars approach navigation with radio metric and optical data

The aerocapture of a space vehicle on hyperbolic approach to Mars results in tight navigation requirements at atmospheric entry. The purpose of this paper is to examine several different methods for approach navigation and to determine what accuracies are possible. The methods are broken into four groups as follows (1) navigation with only Deep Space Network (DSN) tracking of the approach vehicle, (2) navigation with the DSN plus ranging between the approach vehicle and spacecraft in orbit about Mars, (3) navigation with the DSN plus optical data involving the Martian moons, and (4) navigation with DSN range data and differenced range data involving the approach spacecraft and orbiters at Mars. If the current modeling errors that affect earth-based radio metric data, such as errors in tracking station locations, the Martian ephemeris, and differences in the quasar and planetary coordinate frames, are improved, then perhaps earth-based tracking could meet the entry error requirements imposed by aerocapture. If not, then the other three options of intervehicular range, optical data, or differenced range provide highly accurate entry knowledge at least twelve hours before entry.

Konopliv, Alex↗

Visualization Quality Assessment

Understanding how inaccuracies in visualizations affect users’ perception and understanding of scientific data is hard. Inaccuracies in visualizations are quite common and could arise from a range of sources such as errors in the original dataset arising from compression artifacts, errors in the capturing device, noise during transmission of the data, effects due to the algorithm being used to convert data to visualization images, images generated from neural networks, and sources we have yet to discover. Many image quality assessment metrics have been developed to quantify image errors. However, these are usually focused on “natural images” rather than visualizations of scientific data. Common image quality assessment metrics (IQAs) include MSE, PSNR, perceptual metrics such SSIM, FSIM as well as perceptual metrics using deep learning approaches. However, a critical part of understanding how errors are perceived by humans, and subsequently developing more accurate quality assessment metrics, is through user evaluation studies. The goal of this software is to develop a visualization quality assessment (VQA) process that will enable the generation of VQAs that can be used to quantify errors in scientific data visualizations. The VQA development process will include software to support user evaluation experimental design, analysis of visualization differences against standard quality metrics, and the ability to develop additional VQA metrics specific to scientific visualization images.

Grosset, Andre↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

Development and Validation of an Empirical Ocean Color Algorithm with Uncertainties: A Case Study with the Particulate Backscattering Coefficient

We explored how algorithm (model) and in situ measurement (observation) uncertainties can effectively be incorporated into empirical ocean color model development and assessment. In this study we focused on methods for deriving the particulate backscattering coefficient at 555 nm, b(bp)(555)/(m). We developed a simple empirical algorithm for deriving b(bp)(555) as a function of a remote sensing reflectance line height (LH) metric. Model training was performed using a high-quality bio-optical dataset that contains coincident in situ measurements of the spectral remote sensing reflectances, R(rs)(λ)/(sr), and the spectral particulate backscattering coefficients, b(bp)(λ). The LH metric used is defined as the magnitude of Rrs(555) relative to a linear baseline drawn between R(rs)(490) and R(rs)(670). Using an independent validation dataset, we compared the skill of the LH-based model with two other models. We used contemporary validation metrics, including bias and mean absolute error (MAE), that were corrected for model and observation uncertainties. The results demonstrated that measurement uncertainties do indeed impact contemporary validation metrics such as mean bias and MAE. Zeta-scores and z-tests for overlapping confidence intervals were also explored as potential methods for assessing model skill.

ocean color↗

Blind Modeling Validation Exercises Using the Horizontal Dry Cask Simulator

The U.S. Department of Energy (DOE) established a need to understand the thermal-hydraulic properties of dry storage systems for commercial spent nuclear fuel (SNF) in response to a shift towards the storage of high-burnup (HBU) fuel (> 45 gigawatt days per metric ton of uranium, or GWd/MTU). This shift raises concerns regarding cladding integrity, which faces increased risk at the higher temperatures within spent fuel assemblies present within HBU fuel compared to low-burnup fuel (≤ 45 GWd/MTU). A dry cask simulator (DCS) was built at Sandia National Laboratories (SNL) in Albuquerque, New Mexico to produce validation-quality data that can be used to test the accuracy of the modeling used to predict cladding temperatures. These temperatures are critical to evaluating cladding integrity throughout the storage cycle of commercial spent nuclear fuel. A model validation exercise was previously carried out for the DCS in a vertical configuration. Lessons learned during the previous validation exercise have been applied to a new, blind study using a horizontal dry cask simulator (HDCS). Three modeling institutions – the Nuclear Regulatory Commission (NRC), Pacific Northwest National Laboratory (PNNL), and Empresa Nacional del Uranio, S.A., S.M.E. (ENUSA) – were granted access to the input parameters from the DCS Handbook, SAND2017-13058R, and results from a limited data set from the horizontal BWR dry cask simulator tests reported in the HDCS update report, SAND2019-11688R. With this information, each institution was tasked to calculate peak cladding temperatures and air mass flow rates for ten HDCS test cases. Axial as well as vertical and horizontal transverse temperature profiles were also calculated. These calculations were done using modeling codes (ANSYS/Fluent, STAR-CCM+, or COBRA-SFS), each with their own unique combination of modeling assumptions and boundary conditions. For this validation study, the ten test cases of the horizontal dry cask simulator were defined by three independent variables – fuel assembly decay heat (0.5 kW, 1 kW, 2.5 W, and 5 kW), internal backfill pressure (100 kPa and 800 kPa), and backfill gas (helium and air). The plots provided in Chapter 3 of this report show the axial, vertical, and horizontal temperature profiles obtained from the dry cask simulator experiments in the horizontal configuration and the corresponding models used to describe the thermal-hydraulic behavior of this system. The tables provided in Chapter 3 illustrate the closeness of fit of the model data to the experiment data through root mean square (RMS) calculations of the error in peak cladding temperatures (PCTs), PCT axial locations, axial temperature profiles, vertical and horizontal temperature profiles at two different axial locations, and air mass flow rates for the ten test cases, normalized by the experimental results. The model results are assigned arbitrary model numbers to retain anonymity. Due to the relatively flat axial temperature profiles, small temperature gradients resulted in large deviations of all models’ PCT axial location from the experimental PCT axial location. When the PCT axial location error is excluded in the calculation of the combined RMS of the normalized errors that considers PCT, the temperature profiles, and the air mass flow rates, the model data fits the experimental data to within 5%. When the vault information is excluded, the model data fits the experimental data to within 2.5%. An error analysis was developed further for one model, using the model and experimental uncertainties in each validation parameter to calculate validation uncertainties. The uncertainties for each parameter were used to define quantifiable validation criteria. For this analysis, the model was considered validated for a given comparison metric if the normalized error in that metric divided by the validation uncertainty was less than or equal to 1. When considering the combined RMS of the normalized errors of all metrics divided by their validation uncertainties, the model was found to have satisfied the criterion for model validation.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Software errors and complexity: An empirical investigation

The distributions and relationships derived from the change data collected during the development of a medium scale satellite software project show that meaningful results can be obtained which allow an insight into software traits and the environment in which it is developed. Modified and new modules were shown to behave similarly. An abstract classification scheme for errors which allows a better understanding of the overall traits of a software project is also shown. Finally, various size and complexity metrics are examined with respect to errors detected within the software yielding some interesting results.

Basili, V. R.↗

Software errors and complexity: An empirical investigation

The distributions and relationships derived from the change data collected during the development of a medium scale satellite software project show that meaningful results can be obtained which allow an insight into software traits and the environment in which it is developed. Modified and new modules were shown to behave similarly. An abstract classification scheme for errors which allows a better understanding of the overall traits of a software project is also shown. Finally, various size and complexity metrics are examined with respect to errors detected within the software yielding some interesting results.

Basili, Victor R.↗

Evaluating Machine Learning-Based MRI Reconstruction Using Digital Image Quality Phantoms

Quantitative and objective evaluation tools are essential for assessing the performance of machine learning (ML)-based magnetic resonance imaging (MRI) reconstruction methods. However, the commonly used fidelity metrics, such as mean squared error (MSE), structural similarity (SSIM), and peak signal-to-noise ratio (PSNR), often fail to capture fundamental and clinically relevant MR image quality aspects. To address this, we propose evaluation of ML-based MRI reconstruction using digital image quality phantoms and automated evaluation methods. Our phantoms are based upon the American College of Radiology (ACR) large physical phantom but created in k-space to simulate their MR images, and they can vary in object size, signal-to-noise ratio, resolution, and image contrast. Our evaluation pipeline incorporates evaluation metrics of geometric accuracy, intensity uniformity, percentage ghosting, sharpness, signal-to-noise ratio, resolution, and low-contrast detectability. We demonstrate the utility of our proposed pipeline by assessing an example ML-based reconstruction model across various training and testing scenarios. The performance results indicate that training data acquired with a lower undersampling factor and coils of larger anatomical coverage yield a better performing model. The comprehensive and standardized pipeline introduced in this study can help to facilitate a better understanding of the performance and guide future development and advancement of ML-based reconstruction algorithms.

47 OTHER INSTRUMENTATION↗

Scalable Programming Workflows for Validation of Quantum Computers

Hybrid quantum-classical workflows have become standard methods for executing variational algorithms and other quantum simulation techniques, which are key applications for noisy intermediate scale quantum (NISQ) computers. Validating these simulations is an important task which helps gauge the progress of quantum computer development, and classical simulation can serve as a tool to this end. Both exact and more scalable approximate methods with quantifiable error bounds can be used in validation tasks where the applicable metrics include the distance from a calculable ground truth, the quality of an error model fit to data, etc. Here we present a library extension that includes methods for validation of quantum simulations based on scalable hybrid workflows executable on high performance computers. We provide examples that use approximate methods based on tensor networks and stabilizer simulators to bound the error of quantum simulations on NISQ hardware.

Nguyen, Thien↗