Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data efficiency and extrapolation trends in neural network interatomic potentials

Abstract Recently, key architectural advances have been proposed for neural network interatomic potentials (NNIPs), such as incorporating message-passing networks, equivariance, or many-body expansion terms. Although modern NNIP models exhibit small differences in test accuracy, this metric is still considered the main target when developing new NNIP architectures. In this work, we show how architectural and optimization choices influence the generalization of NNIPs, revealing trends in molecular dynamics (MD) stability, data efficiency, and loss landscapes. Using the 3BPA dataset, we uncover trends in NNIP errors and robustness to noise, showing these metrics are insufficient to predict MD stability in the high-accuracy regime. With a large-scale study on NequIP, MACE, and their optimizers, we show that our metric of loss entropy predicts out-of-distribution error and data efficiency despite being computed only on the training set. This work provides a deep learning justification for probing extrapolation and can inform the development of next-generation NNIPs.

36 MATERIALS SCIENCE↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗

Uncertain of uncertainties? A comparison of uncertainty quantification metrics for chemical data sets

Abstract With the increasingly more important role of machine learning (ML) models in chemical research, the need for putting a level of confidence to the model predictions naturally arises. Several methods for obtaining uncertainty estimates have been proposed in recent years but consensus on the evaluation of these have yet to be established and different studies on uncertainties generally uses different metrics to evaluate them. We compare three of the most popular validation metrics (Spearman’s rank correlation coefficient, the negative log likelihood (NLL) and the miscalibration area) to the error-based calibration introduced by Levi et al. ( Sensors 2022 , 22 , 5540). Importantly, metrics such as the negative log likelihood (NLL) and Spearman’s rank correlation coefficient bear little information in themselves. We therefore introduce reference values obtained through errors simulated directly from the uncertainty distribution. The different metrics target different properties and we show how to interpret them, but we generally find the best overall validation to be done based on the error-based calibration plot introduced by Levi et al. Finally, we illustrate the sensitivity of ranking-based methods (e.g. Spearman’s rank correlation coefficient) towards test set design by using the same toy model ferent test sets and obtaining vastly different metrics (0.05 vs. 0.65).

Rasmussen, Maria H.↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗

ESM data downscaling: a comparison of super-resolution deep learning models

Abstract Climate projections at fine spatial resolutions are required to conduct accurate risk assessment for critical infrastructure and design adaptation planning. Generating these projections using advanced Earth system models (ESM) requires significant computational resources. To address this issue, various statistical downscaling techniques have been introduced to generate fine-resolution data from coarse-resolution simulations. In this study, we evaluate and compare five deep learning-based downscaling techniques, namely, super-resolution convolutional neural networks, fast super-resolution convolutional neural network ESM, efficient sub-pixel convolutional neural network, enhanced deep residual network (EDRN), and super-resolution generative adversarial network (SRGAN). These techniques are applied to a dataset generated by the Energy Exascale Earth System Model (E3SM), focusing on key surface variables such as surface temperature, shortwave heat flux, and longwave heat flux. Models are trained and validated using paired fine-resolution (0.25 $$^{\circ }$$ ∘ ) and coarse-resolution (1 $$^{\circ }$$ ∘ ) monthly data obtained from a 9-year simulation. Next, blind testing is performed using monthly data obtained from two different years outside of the training and validation set. To evaluate the efficiency of each technique, different statistical metrics are used, including mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The results show that EDRN outperforms other algorithms in terms of PSNR, SSIM, and MSE, but struggles to capture fine-scale features in the data. In contrast, SRGAN, a generative model that uses perceptual loss, excels in capturing fine details at boundaries and internal structures, resulting in lower LPIPS than other methods.

Pawar, Nikhil M. (ORCID:0000000211613289)↗

Multimodal sensor fusion framework for residential building occupancy detection

For several years now, smart building energy systems have been a research area of intensive activity. In light of the increasing need for sustainable buildings and energy systems, this trend motivates an increasing need for a solution to reduce carbon dioxide emissions and improve energy efficiency. This work proposes a high-performing and transferable occupancy detection framework that combines sensor data from different data modalities, including time series environmental data (temperature, humidity, and illuminance), image data, and acoustic energy data using ensemble method. To draw out the best prediction performance in each modality, the proposed framework was developed, including various models that were designed to learn the occupancy patterns reflected in the physical data streams. To tackle the time series environmental data, we designed two variants of an occupancy detection spatiotemporal pattern network (Occ-STPN) that performs both feature level and decision level fusion, respectively. We also propose a new metric; the fading memory mean square error (FMMSE), that provides a fair evaluation and penalization of delayed occupancy predictions. Multiple open-sourced datasets, including the Electricity Consumption and Occupancy and the University of California, Irvine's (UCI) building occupancy detection dataset, along with our own real data collected from six different houses, were used to validate the algorithms' performance. The experimental results presented herein break down the performance for each sensing modality, and a detailed analysis of the performance is also discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Thermal and electric multidomain dynamic model for integration of power grid distribution with behind-the-meter devices

As renewable energy sources like solar and wind power become more integrated into the grid, coordinated control of behind-the-meter devices is crucial for enhancing grid flexibility and reliability and for meeting cost targets, with standardized models being developed to support this transition. The increasing flexibility and uncertainty of integrated renewable energy grids, along with interactions between various subsystems, make traditional steady-state modeling insufficient to capture transient and dynamic behaviors. Current models (e.g., composite load and battery equivalent models) focus on thermodynamic or electrical characteristics but overlook critical electromechanical interactions. This limits the ability to share performance information for grid services and hampers fast dynamic simulations. In addition, motor stalling is usually triggered by a fault event and attributed to the characteristics of the mechanical torque of the motor, resulting in absorption of a large amount of reactive power during the stalling period. Further, this significant withdrawal of reactive power will deteriorate the dynamic voltage stability of power grids and cause delayed voltage recovery. Therefore, an in-depth modeling of the thermodynamics or mechanical torque is essential to study the impacts of the realistic torque characteristics of those behind-the-meter devices on power system voltage stability. This study developed a dynamic multidomain model for building HVAC systems, such as air-source heat pumps, to simulate their thermal and electrical responses to grid transients. The model can accurately predict power metrics with a mean absolute percentage error of 10%, by validating against with power system computer-aided design performance data. Case studies demonstrate the model capability of capturing the transient response to sudden voltage changes, rapid load fluctuations, and system shutdowns respectively. During a sudden voltage drop (30% for 0.1s), a fully loaded heat pump’s motor speed dropped, continued declining, and shut down after 3.6s, with severe power oscillations and a torque spike. A partially loaded unit experienced temporary oscillations but stabilized. Under higher building loads, compressor speed increased from 64% to 100%, with power and torque rising before stabilizing. In safety-triggered shutdowns, power decreased after minor fluctuations, and torque briefly spiked before dropping to zero.

24 POWER TRANSMISSION AND DISTRIBUTION↗

INTERFACES. A Program for Determining the 3D Structures of Surfaces Sites Using NMR Data

Dynamic nuclear polarization surface enhanced NMR spectroscopy has enabled the determination of high-resolution structures from surface-supported molecules, including singlesite heterogeneous catalysts. Structure determinations have largely mimicked the approaches used in biomolecular NMR spectroscopy, namely, using distance measurements to constrain a conformational search. These early demonstrations made use of purpose-built software, which has limited the adoption of the technique. Herein, we describe the open-source program INTERFACES (Interpret NMR to Elucidate or Reconstruct the Full Atomistic Configurations of External Surfaces) which automates the analysis of RE(SP)DOR data as well as the structure determination for surface sites. Distances, angles, dihedral angles, complex orientation, and distance from the support can all be sampled to find all structures that agree with the experimental data. A χ 2 metric is used to define the error ranges of the REDOR fits and produce structures with an arbitrary level of confidence. Structural solutions are then provided as both overlays and ORTEP-like probability ellipsoids.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Cryogenic-Refined MOSFET Modeling for Oscillator, Frequency Divider, and Amplifier Designs Below 4 K

Capturing device characteristic changes at cryogenic temperatures is crucial for cryo-CMOS circuit designs. In this work, we present an isothermal cryogenic-refined modeling approach for CMOS transistors that is simple, low overhead, and easy to implement while offering the required accuracy for predicting circuit performance at the designated temperatures. Guided by die-level measurement data and circuit design principles, the model introduces corrections to only five critical parameters: threshold voltage, carrier mobility, elevated low-frequency flicker noise, dominant high-frequency shot noise, and subthreshold swing (SS). These refinements are implemented around the foundry-provided SPICE model, which is typically validated only down to about 200 K. With these adjustments, the proposed cryogenic-refined model achieves less than 5% error in both large-signal metrics (I–V characteristics) and small-signal parameters (e.g., transconductance) when compared with device measurements at deep-cryogenic temperatures. The methodology is validated in two advanced technologies: TSMC 40-nm CMOS and GlobalFoundries (GF) 45-nm RF-SOI. We further demonstrate its applicability in three representative RF circuits: a 30-GHz LC oscillator, a high-speed current-mode-logic (CML) frequency divider (FD), and a subthreshold Gb/s amplifier, all showing close agreement between simulated predictions and measurements performed at 4 and 2.5 K. Finally, we believe that the proposed approach is implementation-friendly and can significantly accelerate the development of cryo-CMOS integrated circuits.

circuit modeling↗

A spatial evaluation of Arctic sea ice and regional limitations in CMIP6 historical simulations

The Arctic sea ice response to a warming climate is assessed in a subset of models participating in Phase 6 of the Coupled Model Intercomparison Project (CMIP6), using several metrics in comparison with satellite observations and results from the Pan-Arctic Ice Ocean Modeling and Assimilation System and the Regional Arctic System Model. Our study examines the historical representation of sea ice extent, volume, and thickness using spatial analysis metrics, such as the integrated ice-edge error, Brier score, and Spatial Probability Score. We find that the CMIP6 multi-model mean captures the mean annual cycle and 1979-2014 sea ice trends remarkably well. However, individual models experience a wide range of uncertainty in the spatial distribution of sea ice when compared against satellite measurements and reanalysis data. Our metrics expose common and individual regional model biases, which sea ice temporal analyses alone do not capture. We identify large ice edge and ice thickness errors in Arctic sub-regions, implying possible model specific limitations in or lack of representation of some key physical processes. We postulate that many of them could be related to the oceanic forcing, especially in the marginal and shelf seas, where seasonal sea ice changes are not adequately simulated. We therefore conclude that an individual model’s ability to represent the observed/reanalysis spatial distribution still remains a challenge. We propose the spatial analysis metrics as useful tools to diagnose model limitations, narrow down possible processes affecting them, and guide future model improvements critical to the representation and projections of Arctic climate change.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Updates to the Regional Seismic Travel Time (RSTT) Model: 2. Path-dependent Travel-time Uncertainty

Abstract The regional seismic travel time (RSTT) model and software were developed to improve travel-time prediction accuracy by accounting for three-dimensional crust and upper mantle structure. Travel-time uncertainty estimates are used in the process of associating seismic phases to events and to accurately calculate location uncertainty bounds (i.e. event location error ellipses). We improve on the current distance-dependent uncertainty parameterization for RSTT using a random effects model to estimate slowness (inverse velocity) uncertainty as a mean squared error for each model parameter. The random effects model separates the error between observed slowness and model predicted slowness into bias and random components. The path-specific travel-time uncertainty is calculated by integrating these mean squared errors along a seismic-phase ray path. We demonstrate that event location error ellipses computed for a 90% coverage ellipse metric (used by the Comprehensive Nuclear-Test-Ban Treaty Organization International Data Centre (IDC)), and using the path-specific travel-time uncertainty approach, are more representative (median 82.5% ellipse percentage) of true location error than error ellipses computed using distance-dependent travel-time uncertainties (median 70.1%). We also demonstrate measurable improvement in location uncertainties using the RSTT method compared to the current station correction approach used at the IDC (median 74.3% coverage ellipse).

58 GEOSCIENCES↗

Evaluation of CMIP6 models in simulating the statistics of extreme precipitation over Eastern Africa

We report the Eastern Africa region experiences frequent extreme precipitation events that can cause destruction of property and environment, and loss of lives. Thus, there is a need to understand how these events may change in the future and how well the global climate models that are used to make projections can simulate precipitation extremes in this region before they can be used in downscaling or flood and drought impact assessment studies. In this work, we evaluated the ability of sixteen Coupled Model Intercomparison Project Phase 6 (CMIP6) models to simulate present-day precipitation extremes over the Eastern Africa region during the two rainy seasons (March–May and September–November). We used nine extreme precipitation indices (including seven (one) indices of wet (dry) extremes) defined by the Expert Team on Climate Change Detection and Indices. The CMIP6 models were evaluated against two gridded observation datasets: Global Precipitation Climatology Project One-Degree Daily Dataset and Tropical Rainfall Measuring Mission Multi-satellite Precipitation Analysis 3B42. Three model performance metrics (percentage bias, normalized root-mean-square error, and pattern correlation coefficient) were employed to further assess the strengths and weakness of the models. Our results show that the multi-model ensemble mean generally provides a better representation of observed precipitation and related extremes compared to individual models when considering all metrics and seasons. Several consistent biases are evident across CMIP6 models, which tend to overestimate the total-wet day precipitation and consecutive wet days, and underestimate very wet days and maximum 5-day precipitation in both seasons. Furthermore, no single model consistently performs best, model performance varies with the season and index under consideration and is generally independent of horizontal resolution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Validation Analysis of Medium-Scale Methanol Pool Fire Simulated in SIERRA/Fuego [Slides]

Analysis of methanol pool fire conducted as part of validation study for SIERRA/Fuego. Results showed trends & errors consistent with related studies. Area validation metric provides way to quantify model form uncertainty. AVM shows that more work could be done to understand how model form uncertainty varies with mesh resolution. There is a possible atypical use of MAVM on time-series data. AVM shows mismatch between predicted flame height and experimental value less sensitive to variations in mixture fraction than temperature. Mismatch about experimental value also more symmetric for mixture fraction. Our analysis showed that mixture fraction is preferable for this application.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Investigating Tropical Versus Extratropical Influences on the Southern Hemisphere Tropical Edge in the Unified Model

Abstract Since the late 1970s, observations have shown a widening of the tropical Hadley cell (HC) circulation. State‐of‐the‐art climate models reproduce the general trend along with a projected continuous expansion. Discrepancies in expansion rates of observation‐ and model‐based studies have been attributed to differences in applied methods, natural variability and model shortcomings. Furthermore, the driving influence of tropical or extratropical processes on these changes is not well understood. All of this highlights the dynamical mechanisms and the region of origin controlling the tropical width are still insufficiently understood. Here we examine the influence of systematic model biases of the atmosphere‐only Unified Model (UM) onto the simulation of the Southern Hemisphere (SH) tropical edge. We utilize nudged experiments with prescribed sea surface temperatures, where potential temperature and horizontal winds are relaxed back to ERA‐Interim reanalysis for a 20‐year period in selected regions. Correcting model biases in the tropics and extratropics separately allows us to dissect the dominant remote impacts of present model errors onto the SH tropical edge simulation. The experiments are applied to established tropical width metrics ranging from near‐surface to upper‐level metrics capturing the poleward flank of the HC. We find both regions work remotely to reduce errors in the UM fields and location of the tropical edge. Surprisingly, correcting the extratropical biases, south of 45°S, more consistently improves the tropical width across the metrics and seasons than nudging the tropics (10°N–10°S). These findings demonstrate the substantial role of extratropical influences in locating the SH tropical edge.

54 ENVIRONMENTAL SCIENCES↗

Toward Guided Mutagenesis: Gaussian Process Regression Predicts MHC Class II Antigen Mutant Binding

Antigen-specific immunotherapies (ASI) require successful loading and presentation of antigen peptides into the major histocompatibility complex (MHC) binding cleft. One route of ASI design is to mutate native antigens for either stronger or weaker binding interaction to MHC. Exploring all possible mutations is costly both experimentally and computationally. To reduce experimental and computational expense, here we investigate the minimal amount of prior data required to accurately predict the relative binding affinity of point mutations for peptide-MHC class II (pMHCII) binding. Using data from different residue subsets, we interpolate pMHCII mutant binding affinities by Gaussian process (GP) regression of residue volume and hydrophobicity. We apply GP regression to an experimental data set from the Immune Epitope Database, and theoretical data sets from NetMHCIIpan and Free Energy Perturbation calculations. We find that GP regression can predict binding affinities of nine neutral residues from a six-residue subset with an average R 2 coefficient of determination value of 0.62 ± 0.04 (±95% CI), average error of 0.09 ± 0.01 kcal/mol (±95% CI), and with an receiver operating characteristic (ROC) AUC value of 0.92 for binary classification of enhanced or diminished binding affinity. Similarly, metrics increase to an R2 value of 0.69 ± 0.04, average error of 0.07 ± 0.01 kcal/mol, and an ROC AUC value of 0.94 for predicting seven neutral residues from an eight-residue subset. Our work finds that prediction is most accurate for neutral residues at anchor residue sites without register shift. This work holds relevance to predicting pMHCII binding and accelerating ASI design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Monitoring agroecosystem productivity and phenology at a national scale: A metric assessment framework

Effective measurement of seasonal variations in the timing and amount of production is critical to managing spatially heterogeneous agroecosystems in a changing climate. Although numerous technologies for such measurements are available, their relationships to one another at a continental extent are unknown. Using data collected from across the Long-Term Agroecosystem Research (LTAR) network and other networks, we investigated correlations among key metrics representing primary production, phenology, and carbon fluxes in croplands, grazing lands, and crop-grazing integrated systems across the continental U.S. Metrics we examined included gross primary productivity (GPP) estimated from eddy covariance (EC) towers and modelled from the Landsat satellite, Landsat NDVI, and vegetation greenness (Green Chromatic Coordinate, GCC) from tower-mounted PhenoCams for 2017 and 2018. Overall, our analysis compared production dynamics estimated from three independent ground and remote platforms using data for 34 agricultural sites constituting 51 site-years of co-located time series. Pairwise sensor comparisons across all four metrics revealed stronger correlation and lower root mean square error (RMSE) between end of season (EOS) dates (Pearson R ranged from 0.6 to 0.7 and RMSE from 32.5 to 67.8) than start of season (SOS) dates (0.46 to 0.69 and 40.4 to 66.2). Overall, moderate to high correlations between SOS and EOS metrics complemented one another except at some lower productivity grazing land sites where estimating SOS can be challenging. Growing season length estimates derived from 16-day satellite GPP (179.1 days) were significantly longer than those from PhenoCam G CC (70.4 days, p adj < 0.0001) and EC GPP (79.6 days, p adj < 0.0001). Landscape heterogeneity did not explain differences in SOS and EOS estimates. Annual integrated estimates of productivity from EC GPP and PhenoCam G CC diverged from those estimated by Landsat GPP and NDVI at sites where annual production exceeds 1000 gC/m –2 yr –1 . Based on our results, we developed a “metric assessment framework” that articulates where and how metrics from satellite, eddy covariance and PhenoCams complement, diverge from, or are redundant with one another. The framework was designed to optimize instrumentation selection for monitoring, modeling, and forecasting ecosystem functioning with the ultimate goal of informing decision-making by land managers, policy-makers, and industry leaders working at multiple scales.

54 ENVIRONMENTAL SCIENCES↗

Evaluating simplifications of subsurface process representations for field-scale permafrost hydrology models

Abstract. Permafrost degradation within a warming climate poses a significant environmental threat through both the permafrost carbon feedback and damage to human communities and infrastructure. Understanding this threat relies on better understanding and numerical representation of thermo-hydrological permafrost processes and the subsequent accurate prediction of permafrost dynamics. All models include simplified assumptions, implying a tradeoff between model complexity and prediction accuracy. The main purpose of this work is to investigate this tradeoff when applying the following commonly made assumptions: (1) assuming equal density of ice and liquid water in frozen soil, (2) neglecting the effect of cryosuction in unsaturated freezing soil, and (3) neglecting advective heat transport during soil freezing and thaw. This study designed a set of 62 numerical experiments using the Advanced Terrestrial Simulator (ATS v1.2) to evaluate the effects of these choices on permafrost hydrological outputs, including both integrated and pointwise quantities. Simulations were conducted under different climate conditions and soil properties from three different sites in both column- and hillslope-scale configurations. Results showed that amongst the three physical assumptions, soil cryosuction is the most crucial yet commonly ignored process. Neglecting cryosuction, on average, can cause 10 %–20 % error in predicting evaporation, 50 %–60 % error in discharge, 10 %–30 % error in thaw depth, and 10 %–30 % error in soil temperature at 1 m beneath the surface. The prediction error for subsurface temperature and water saturation is more obvious at hillslope scales due to the presence of lateral flux. By comparison, using equal ice–liquid density has a minor impact on most hydrological metrics of interest but significantly affects soil water saturation with an averaged 5 %–15 % error. Neglecting advective heat transport presents the least error, 5 % or even much lower, in most metrics of interest for a large-scale Arctic tundra system without apparent influence caused by localized groundwater flow, and it can decrease the simulation time at hillslope scales by 40 %–80 %. By challenging these commonly made assumptions, this work provides permafrost hydrology scientists an important context for understanding the underlying physical processes, including allowing modelers to better choose the appropriate process representation for a given modeling experiment.

54 ENVIRONMENTAL SCIENCES↗