Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Effects of machine learning errors on human decision-making: manipulations of model accuracy, error types, and error importance

Abstract This study addressed the cognitive impacts of providing correct and incorrect machine learning (ML) outputs in support of an object detection task. The study consisted of five experiments that manipulated the accuracy and importance of mock ML outputs. In each of the experiments, participants were given the T and L task with T-shaped targets and L-shaped distractors. They were tasked with categorizing each image as target present or target absent. In Experiment 1, they performed this task without the aid of ML outputs. In Experiments 2–5, they were shown images with bounding boxes, representing the output of an ML model. The outputs could be correct (hits and correct rejections), or they could be erroneous (false alarms and misses). Experiment 2 manipulated the overall accuracy of these mock ML outputs. Experiment 3 manipulated the proportion of different types of errors. Experiments 4 and 5 manipulated the importance of specific types of stimuli or model errors, as well as the framing of the task in terms of human or model performance. These experiments showed that model misses were consistently harder for participants to detect than model false alarms. In general, as the model’s performance increased, human performance increased as well, but in many cases the participants were more likely to overlook model errors when the model had high accuracy overall. Warning participants to be on the lookout for specific types of model errors had very little impact on their performance. Overall, our results emphasize the importance of considering human cognition when determining what level of model performance and types of model errors are acceptable for a given task.

97 MATHEMATICS AND COMPUTING↗

Information transmission with continuous variable quantum erasure channels

Quantum capacity, as the key figure of merit for a given quantum channel, upper bounds the channel's ability in transmitting quantum information. Identifying different types of channels, evaluating the corresponding quantum capacity, and finding the capacity-approaching coding scheme are the major tasks in quantum communication theory. Quantum channel in discrete variables has been discussed enormously based on various error models, while error model in the continuous variable channel has been less studied due to the infinite dimensional problem. In this paper, we investigate a general continuous variable quantum erasure channel. By defining an effective subspace of the continuous variable system, we find a continuous variable random coding model. We then derive the quantum capacity of the continuous variable erasure channel in the framework of decoupling theory. The discussion in this paper fills the gap of a quantum erasure channel in continuous variable setting and sheds light on the understanding of other types of continuous variable quantum channels.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Model Form Error Correction for a Black-Box Thermal Battery Heat Transfer Simulation

Thermal batteries are crucial for supplying power to high-consequence engineering applications such as rockets. Computational simulations have been developed to predict thermal battery behavior, but these simulations often suffer from modeling errors, including model form uncertainty. Addressing this uncertainty can be achieved by quantifying either the model discrepancy in the output or the model form error (MFE) in the governing equation. MFE is particularly valuable as it can be better extrapolated beyond observed outputs, which is essential for predictions involving changes in external system loading, system configuration and geometry, or output quantities. This paper employs a state estimation approach to estimate MFE using experimental data and then utilizes machine learning (ML) to model its relationship with state variables. A nonintrusive technique is used to estimate MFE in a black-box thermal battery heat transfer simulation. The trained machine learning model for MFE is then applied to correct simulation predictions under extrapolated initial conditions and battery configurations. In conclusion, the methodology's performance is evaluated using additional experimental data, demonstrating its effectiveness in improving prediction accuracy.

Batteries↗

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases

A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.

97 MATHEMATICS AND COMPUTING↗

Characterization of Partially Observed Epidemics - Application to COVID-19

This report documents a statistical method for the "real-time" characterization of partially observed epidemics. Observations consist of daily counts of symptomatic patients, diagnosed with the disease. Characterization, in this context, refers to estimation of epidemiological parameters that can be used to provide short-term forecasts of the ongoing epidemic, as well as to provide gross information for the time-dependent infection rate. The characterization problem is formulated as a Bayesian inverse problem, and is predicated on a model for the distribution of the incubation period. The model parameters are estimated as distributions using a Markov Chain Monte Carlo (MCMC) method, thus quantifying the uncertainty in the estimates. The method is applied to the COVID-19 pandemic of 2020, using data at the country, provincial (e.g., states) and regional (e.g. county) levels. The epidemiological model includes a stochastic component due to uncertainties in the incubation period. This model-form uncertainty is accommodated by a pseudo-marginal Metropolis-Hastings MCMC sampler, which produces posterior distributions that reflect this uncertainty. We approximate the discrepancy between the data and the epidemiological model using Gaussian and negative binomial error models; the latter was motivated by the over-dispersed count data. For small daily counts we find the performance of the calibrated models to be similar for the two error models. For large daily counts the negative-binomial approximation is numerically unstable unlike the Gaussian error model. Application of the model at the country level (for the United States, Germany, Italy, etc.) generally provided accurate forecasts, as the data consisted of large counts which suppressed the day-to-day variations in the observations. Further, the bulk of the data is sourced over the duration before the relaxation of the curbs on population mixing, and is not confounded by any discernible country-wide second wave of infections. At the state-level, where reporting was poor or which evinced few infections (e.g., New Mexico), the variance in the data posed some, though not insurmountable, difficulties, and forecasts were able to capture the data with large uncertainty bounds. The method was found to be sufficiently sensitive to discern the flattening of the infection and epidemic curve due to shelter-in-place orders after around 90% quantile for the incubation distribution (about 10 days for COVID-19). The proposed model was also used at a regional level to compare the forecasts for the central and north-west regions of New Mexico. Modeling the data for these regions illustrated different disease spread dynamics captured by the model. While in the central region the daily counts peaked in the late April, in the north-west region the ramp-up continued for approximately three more weeks.

59 BASIC BIOLOGICAL SCIENCES↗

Improved Understanding of Coupled Water and Carbon Cycle Processes through Machine Learning Approaches

Focal Area(s): This white paper addresses how explainable machine learning (ML) algorithms can improve insights gained from complex data (Focal Area 3). We will also address how ML approaches related to sensor compression and low energy AI hardware can be used for efficient data acquisition (Focal Area 1). Science Challenge: We will focus on the coupled water cycle and carbon cycle processes in terrestrial ecosystems and terrestrial-aquatic interfaces. Thus, our approaches will be centered around several data-model integration challenges indicated in EESD strategic plans for two science focus areas: Terrestrial Ecosystem Science and Hydrobiogeochemistry. We can use AI to correct systematic model errors due to biases in either data or model structure/processes. Error patterns in model predictions usually vary by region, model, season etc. But there are systematic patterns in them. e.g., soil moisture tends to be overestimated in the arid western continental United States underestimated in wetter eastern USA; some land surface models tend to underestimate moisture in wet seasons and overestimate in dry seasons. Moreover, the consequences of extreme events (e.g., drought, extreme flooding, storm surges associated with tropical storms, hurricanes) on carbon cycle processes are not well represented in ecosystem and Earth system models. Redox-sensitive processes (e.g., rapid oxidation/reduction of iron and other redox-sensitive elements in soil microsites subjected to fluctuated hydrology) in terrestrial-aquatic interfaces further challenge model predictions of hot-spots (and hot-moments) due to poor understanding of underlying mechanisms. Thus, AI/ML approaches can be used to learn patterns in the data and model errors and use them to build model equations and correct process-based model errors.

54 ENVIRONMENTAL SCIENCES↗

PyOED: An Extensible Suite for Data Assimilation and Model-Constrained Optimal Design of Experiments

This article describes PyOED, a highly extensible scientific package that enables developing and testing model-constrained optimal experimental design (OED) for inverse problems. Specifically, PyOED aims to be a comprehensive Python toolkit for model-constrained OED. The package targets scientists and researchers interested in understanding the details of OED formulations and approaches. It is also meant to enable researchers to experiment with standard and innovative OED technologies with a wide range of test problems (e.g., simulation models). OED, inverse problems (e.g., Bayesian inversion), and data assimilation (DA) are closely related research fields, and their formulations overlap significantly. Thus, PyOED is continuously being expanded with a plethora of Bayesian inversion, DA, and OED methods as well as new scientific simulation models, observation error models, and observation operators. These pieces are added such that they can be permuted to enable testing OED methods in various settings of varying complexities. The PyOED core is completely written in Python and utilizes the inherent object-oriented capabilities; however, the current version of PyOED is meant to be extensible rather than scalable. Specifically, PyOED is developed to “enable rapid development and benchmarking of OED methods with minimal coding effort and to maximize code reutilization.” This article provides a brief description of the PyOED layout and philosophy and provides a set of exemplary test cases and tutorials to demonstrate the potential of the package.

97 MATHEMATICS AND COMPUTING↗

MultiPEM Toolbox: User Manual [Rev. 2]

This document explains use of the Multi-Phenomenology Explosion Monitoring (Multi PEM) Toolbox, a collection of R scripts for estimating the unknown device parameters of a new event with uncertainty quantification. The methodology and application used for illustration in this user manual are fully documented in a Los Alamos National Laboratory technical report hereafter designated “WPA” for reference. Additional details on the application are found in a recent journal article. Two assessment types are available: rapid and complete. Rapid assessments are conducted in two stages, as described in Section 2. In the first stage, calibration data are used to estimate forward and error model parameters (WPA, §5.1) and (if relevant) errors-in-variables yield values for calibration sources (WPA, §3, Equation (3)). In the second stage, new event data are used to estimate the unknown new event device parameters (WPA, §5.2) with uncertainty quantification. Two options for treating the inferred first stage parameters in second stage Bayesian analysis are available: fixing them at their maximum likelihood estimate (default), or multiple imputation. Multiple imputation involves utilizing several posterior samples (imputations) of the first stage parameters as fixed values in the second stage posterior sampling of the new event device parameters. Second stage sampling is conducted across imputations in parallel to improve computational efficiency. This method produces improved uncertainty quantification of the new event device parameters compared with the default treatment of the first stage parameters, at the expense of additional computation. Complete assessments are conducted in a single stage, as described in Section 3. Calibration and (if relevant) new event data are used simultaneously to estimate all forward model, error model, and (if relevant) new event device parameters with uncertainty quantification on the latter. As the name suggests, rapid assessments generally run substantially faster than complete assessments (even with multiple imputation), because the results of first stage analysis can be stored and incorporated into estimating a relatively low-dimensional space of new event device parameters whenever relevant new event data becomes available. On the other hand, complete assessments must be run on the full set of model and device parameters with calibration and new event data every time the latter becomes available.

97 MATHEMATICS AND COMPUTING↗

Model-form Error Correction using Universal Differential Equations for an Agent-Based Model of Infectious Disease

This report demonstrates universal differential equations (UDEs) as an approach to bridge the gap between ordinary differential equations (ODE) models and agent-based models (ABMs). Using UDE models as surrogates for ABMs allows us to preserve the foundational ODE that represents global disease dynamics while coupling it with a neural network model to approximate functions for the local behaviors of the ABM.

59 BASIC BIOLOGICAL SCIENCES↗

Accounting for erroneous model structures in biokinetic process models

In engineering practice, model-based design requires not only a good process-based model, but also a good description of stochastic disturbances and measurement errors to learn credible parameter values from observations. However, typical methods use Gaussian error models, which often cannot describe the complex temporal patterns of residuals. Consequently, this results in overconfidence in the identified parameters and, in turn, optimistic reactor designs. Here, we assess the strengths and weaknesses of a method to statistically describe these patterns with autocorrelated error models. This method produces increased widths of the credible prediction intervals following the inclusion of the bias term, in turn leading to more conservative design choices. However, we also show that the augmented error model is not a universal tool, as its application cannot guarantee the desired reliability of the resulting wastewater reactor design.

42 ENGINEERING↗

Model uncertainty in accelerator application simulations

Monte-Carlo nuclear reaction and transport codes are widely used to devise accelerator-based nuclear physics experiments; at the same time, many experiments are performed to validate the Monte-Carlo codes, which can be used for the design of full-scale nuclear power applications or the design of new benchmark experiments. Dedicated model benchmark studies investigate a broad range of nuclear reactions and quantities. Examples of these include isotope formation or secondary particle fluxes that result from the interactions of GeV-range hadrons with monoisotopic targets, which can be used to assess the respective systematic uncertainty of models. Such benchmark studies, as well as many nuclear application experiments and simulations carried out by various groups over the last few decades, enable us to draw methodological lessons. In this work, model uncertainty determined based on available experimental data allow us to identify the effects of practitioner expertise as well as the design of codes (user access to micro-scale parameters) on the range of uncertainties. We found that in cases when simulations are performed by code developers or users that are very experienced in performing simulations, the model to experiment quantity ratios generally agree with the limits determined by dedicated benchmark studies. In other cases, the ratios generally tend to be either smaller (underestimation of model error) or larger (overestimation of model error). A plausible explanation of the aforementioned effects is suggested.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Interpolation Models and Error Bounds for Verifiable Scientific Machine Learning

This repository contains python scripts and numerical data accompanying the paper: "Leveraging Interpolation Models and Error Bounds for Verifiable Scientific Machine Learning," Tyler Chang, Andrew Gillette, Romit Maulik, 2024. The following subdirectories are included: - "interpolants" contains our interpolation scripts used for all studies - "experiments" contains scripts demonstrating our experiments with synthetic data - "airfoil" contains scripts demonstrating our experiments with the publicly available UIUC airfoil dataset. Further instructions are provided in READMEs within the sub-directories.

Gillette, Andrew↗

Quantifying modeling uncertainty in simplified beam models for building response prediction

The use of simple models for response prediction of building structures is preferred in earthquake engineering for risk evaluations at regional scales, as they make computational studies more feasible. The primary impediment in their gainful use presently is the lack of viable methods for quantifying (and reducing upon) the modeling errors/uncertainties they bear. This study presents a Bayesian calibration method wherein the modeling error is embedded into the parameters of the model. Here, the method is specifically described for coupled shear-flexural beam models here, but it can be applied to any parametric surrogate model. The major benefit the method offers is the ability to consider the modeling uncertainty in the forward prediction of any degree-of-freedom or composite response regardless of the data used in calibration. The method is extensively verified using two synthetic examples. In the first example, the beam model is calibrated to represent a similar beam model but with enforced modeling errors. In the second example, the beam model is used to represent the detailed finite element model of a 52-story building. Both examples show the capability of the proposed solution to provide realistic uncertainty estimation around the mean prediction.

47 OTHER INSTRUMENTATION↗

Successive Procedure for Solution Verification Based on User Needs

This paper discusses a revised solution verification procedure for computational fluid dynamics simulations to estimate the uncertainties in the quantities of interest based on discretization error models. This proposed procedure builds upon current procedures described in ASME V&V 20 but provides more guidance in determining the necessary number of mesh levels to build reliable discretization error models. Such guidance is particularly useful for practicing engineers without prior experience in solution verification. The key features of this proposed solution verification procedure are the ability to determine the need for additional mesh levels iteratively and the seamless treatment for underdetermined, exact, and overdetermined solutions of the power series approximation to the discretization error models. This study applies the proposed procedure to a set of synthetic examples to demonstrate the revised procedure’s clarity in determining the number of mesh solutions required for a reliable estimate of the discretization error in computational fluid dynamics settings. Additionally, this proposed procedure prevents a potential pathway in the current procedure in ASME V&V 20 that may lead to unreasonably small discretization errors.

Weinmeister, Justin↗

Quantification of modeling uncertainty in the Rayleigh damping model

Understanding and accurately characterizing energy dissipation mechanisms in civil structures during earthquakes is an important element of seismic assessment and design. The most commonly used model is attributed to Rayleigh. This paper proposes a systematic approach to quantify the uncertainty associated with Rayleigh's damping model. Bayesian calibration with embedded model error is employed to treat the coefficients of the Rayleigh model as random variables using modal damping ratios. Through a numerical example, we illustrate how this approach works and how the calibrated model can address modeling uncertainty associated with the Rayleigh damping model.

42 ENGINEERING↗

A novel network-based approach to determining measurement representation error for model evaluation of aerosol microphysical properties

Atmospheric aerosol size and abundance influence radiative effects and climate change. To date, efforts to constrain global climate models’ radiative forcing with in situ aerosol observations have been hamstrung by uncertainty. One source of error, the regional “representation error,” arises when accurate but sparse single-point measurements of atmospheric aerosol distributions are compared with a model value, assuming that the single-point measurement is representative of the model domain. The Portable Optical Particle Spectrometer network in the Southern Great Plains (POPSnet-SGP) campaign has demonstrated that a network of nearly autonomous aerosol instruments operating at ambient temperature and relative humidity (with low measurement error) may be used to quantify measurement representation error and investigate the factors introducing heterogeneity in aerosol distributions across a rural, continental background region. Measurements were made using Portable Optical Particle Spectrometer (POPS) instruments at several sites for five months across the Department of Energy’s Aerosol Radiation Measurement Southern Great Plains (ARM-SGP) User Facility in the central USA. Measurement representation error decreased with longer averaging periods (20-40 % between 1 sec and 1 day), varied between sites by 10 – 20 % for aerosol concentration 140 – 2500 nm in diameter (N_140), and was higher for aerosols > 400 nm in diameter (N_400). Our measurements also show the influence of local meteorology on aerosol surface area (A_140) and size distributions: A_140 is positively correlated with wind speed and relative humidity, negatively correlated with precipitation, and lower given westerly winds. Based on this study, we conclude that the POPSnet approach provides considerably more insight into the spatial variability in the aerosol population that can be used to constrain climate models than would be available from similar networks of PM 2.5 monitors.

54 ENVIRONMENTAL SCIENCES↗