Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Power Flow Geometry and Approximation

Here, the power flow equations are important in numerous power systems problems of practical interest which consider alternating current power flow (ACPF) physics. Perhaps the most well studied being the alternating current optimal power flow problem (ACOPF), seeking to optimize the operation of an electric power system. Due to their non-linearity, problems which include the power flow equations are typically challenging, particularly in optimization. Interestingly, the set of solutions to the power flow equations forms a smooth manifold. As a result, differential geometry can be used to describe and analyze this set of equations. This approach has proven effective in several engineering applications (e.g., solving ACOPF and analyzing the solution space boundary). Central to the success of this approach is an understanding of the power flow manifold's geometry. In this work, we develop the geometric and topological properties of this manifold using concepts from differential geometry. After demonstrating the convenience of this manifold's representation as a function's graph, computational methods are emphasized: we develop retractions, error bounds for linear approximation, and formulas for evaluating the Riemannian metric (including associated objects such as geodesics and the curvature tensor). Scalar curvature and the second fundamental form play a new role in quantifying the quality of linear approximations, like the popular direct current approximation. All functions are implemented in Julia and available in an online repository. Proofs are included for completeness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

On the use of air temperature and precipitation as surrogate predictors in soil respiration modelling

Soil respiration (R S ), the soil-to-atmosphere CO 2 flux that is a major component of the global carbon cycle, is strongly influenced by local soil temperature (T soil ) and water content (SWC). Regional to global-scale R S modelling thus requires this information at local scales, but few high-quality, wall-to-wall (global) T soil and SWC data exist. As a result, such modelling efforts commonly use air temperature (T air ) and monthly precipitation (P m ) as surrogate predictors, but their site-scale accuracy and potential bias are unknown. In this report we used monthly data from 880 sites across a wide variety of different environmental conditions (i.e., climate, ecosystem type, elevation, vegetation leaf habit and drainage conditions) to determine the suitability of T air as a surrogate for T soil , and data from 507 sites to examine the suitability of P m as a surrogate for SWC. Site-specific linear and second-order exponential non-linear models were compared using model evaluation metrics (i.e., slope, p-value of slope, root mean square error [RMSE], index of agreement and model efficiency). We found that T soil and T air are highly correlated and explain similar R S variability. In contrast, P m is not a good surrogate for SWC, even though P m explains a similar amount of R S variability to SWC. The wide variability in the site-specific relationships between R S and SWC means that no single relationship can be used for large-scale modelling. The results from this study support the use of T air in continental-to-global scale R S models, and highlight the urgent need for continental-to-global scale SWC datasets for the modelling and evaluation of future soil carbon dynamics under global climate change.

54 ENVIRONMENTAL SCIENCES↗

Phase-based velocity extraction method for photonic Doppler velocimetry with potential higher time resolution

We present an extension of the [Takeda et al., J. Opt. Soc. Am. 72, 156 (1982)] phase extraction method to heterodyne photonic Doppler velocimetry applications. The method yields results equivalent to those obtained by the short-time Fourier transform (STFT), while offering potential improvements in time resolution. Unlike STFT, which relies on window functions, such as the Hamming window, that emphasize central data points and diminish the influence of edges, the extended Takeda method utilizes all data uniformly. This uniform treatment allows for the derivation of empirical equations that directly relate velocity error to the actual time resolution rather than to the local analysis duration. The established equation provides a useful metric for both optimizing hardware configuration and guiding data analysis. Simulation and experimental results confirm that, for a given dataset, specifying a target time resolution yields consistent velocity errors for both methods. These findings underscore the Takeda method’s advantages, particularly its potential higher time resolution and reduced computational burden, making it a valuable tool for high-throughput applications such as laser dynamic compression experiments.

Computer simulation↗

Model validation and error attribution for a drifting qubit

Qubit performance is often reported in terms of a variety of single-value metrics, each providing a facet of the underlying noise mechanism limiting performance. However, the value of these metrics may drift over long timescales, and reporting a single number for qubit performance fails to account for the low-frequency noise processes that give rise to this drift. Here, in this work, we demonstrate how we can use the distribution of these values to validate or invalidate candidate noise models. We focus on the case of randomized benchmarking (RB), where typically a single error rate is reported but this error rate can drift over time when multiple passes of RB are performed. We show that using a statistical test as simple as the Kolmogorov-Smirnov statistic on the distribution of RB error rates can be used to rule out noise models, assuming the experiment is performed over a long enough time interval to capture relevant low frequency noise. With confidence in a noise model, we show how care must be exercised when performing error attribution using the distribution of drifting RB error rate.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

Fast baryonic field painting for Sunyaev-Zel’dovich analyses: Transfer function vs hybrid effective field theory

Here, we present two approaches for “painting” baryonic properties relevant to the Sunyaev-Zel’dovich (SZ) effect—optical depth and Compton-y—onto three-dimensional N-body simulations, using the MillenniumTNG suite as a benchmark. The goal of these methods is to produce fast and accurate reconstruction methods to aid future analyses of baryonic feedback using the SZ effect. The first approach employs a Gaussian process emulator to model the SZ quantities via a transfer function, while the second utilizes hybrid effective field theory (HEFT) to reproduce these quantities within the simulation. Our analysis involves comparing both methods to the true MillenniumTNG optical depth and Compton-y fields using several metrics, including the cross-correlation coefficient, power spectrum, and power spectrum error. Additionally, we assess how well the reconstructed fields correlate with dark matter haloes across various mass thresholds. The results indicate that the transfer function method yields more accurate reconstructions for fields with initially high correlations (r ≈ 1), such as between the optical depth and dark matter fields. Conversely, the HEFT-based approach proves more effective in enhancing correlations for fields with weaker initial correlations (r ∼ 0.5), such as between the Compton-y and dark matter fields. Lastly, we discuss extensions of our methods to improve the reconstruction performance at the field level.

Liu, R. Henry [University of California, Berkeley,↗

Towards Full Field-of-View Fourier Ptychography for Extreme Ultraviolet Microscope

We evaluate various Fourier ptychographic microscopy (FPM) reconstruction algorithms using both simulated and experimental data acquired from an Extreme Ultraviolet (EUV, 13.5 nm wavelength) microscope. We specifically focus on the algorithms' ability to robustly address field-dependent aberrations, which enables increased spatial resolution and quantitative phase imaging across an expanded field of view. We systematically compare the algorithms' performance under aberrations for a single zoneplate imaging system, utilizing Fourier Ring Correlation (FRC) as a systematic metric for assessing reconstruction quality. Furthermore, we explore the impact of systematic errors on the reconstruction of experimental data, aiming to increase the effective field of view by 25-fold, from the nominal 5x5 um2 diffraction-limited area. Additionally, our evaluation incorporates innovative FPM-adjacent methodologies, including the Angular Ptychographic Imaging with Closed-form method (APIC), for reconstructing EUV images.

Gu, Chaoying↗

Trust Model System for the Energy Grid of Things Network Communications

Network communication is crucial in the Energy Grid of Things (EGoT). Without a network connection, the energy grid becomes just a power grid where the energy resources are available to the customer uni-directionally. A mechanism to analyze and optimize the energy usage of the grid can only happen through a medium, a communications network, that enables information exchange between the grid participants and the service provider. Security implementers of EGoT network communication take extraordinary measures to ensure the safety of the energy grid, a critical infrastructure, as well as the safety and privacy of the grid participants. With the dynamic nature of network communication of the EGoT, the information provided by the customer or the service provider can be falsified by a malicious attacker. Therefore, a trust model is necessary to monitor any abnormal activities. This paper describes a distributed trust model system that meets the need of the EGoT. This paper describes methods for evaluating and improving the distributed trust model using standard hypothesis testing metrics such as true positive, false positive, true negative, false negative, equal error rate, and F1 score. Example calculations are shown based on generated sample data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

Development and evaluation of a new 4DEnVar-based weakly coupled ocean data assimilation system in E3SMv2

The development, implementation, and evaluation of a new weakly coupled ocean data assimilation (WCODA) system for the fully coupled Energy Exascale Earth System Model version 2 (E3SMv2) utilizing the four-dimensional ensemble variational (4DEnVar) method are presented in this study. The 4DEnVar method, based on the dimension-reduced projection four-dimensional variational (DRP-4DVar) approach, replaces the adjoint model with the ensemble technique, thereby reducing computational demands. Monthly mean ocean temperature and salinity data from the EN4.2.1 reanalysis are integrated into the ocean component of E3SMv2 from 1950 to 2021 with the goal of providing realistic initial conditions for decadal predictions and predictability studies. The performance of the WCODA system is assessed using various metrics, including the reduction rate of the cost function, root mean square error (RMSE) differences, correlation differences, and model biases. Results indicate that the WCODA system effectively assimilates the reanalysis data into the climate model, consistently achieving negative reduction rates of the cost function and notable improvements in RMSE and correlation across various ocean layers and regions. Significant enhancements are observed in the upper ocean layers across the majority of global ocean regions, particularly in the north Atlantic, north Pacific, and Indian Ocean. Model biases in sea surface temperature and salinity are also substantially reduced. For sea surface temperature, cold biases in the north Pacific and north Atlantic are diminished by about 1–2 °C, and warm biases in the Southern Ocean are corrected by approximately 1.5–2.5 °C. In terms of salinity, improvements are observed with bias reductions of about 0.5–1 psu in the north Atlantic and north Pacific and up to 1.5 psu in parts of the Southern Ocean. The ultimate goal of the WCODA system is to advance the predictive capabilities of E3SM for subseasonal to decadal climate predictions, thereby supporting research on strategic energy-sector policies and planning.

54 ENVIRONMENTAL SCIENCES↗

Count Every Trip: Finding the Uncertainty in Energy Estimates Made from Inferred Travel Modes

To properly inform transport policy and infrastructure changes, transportation related metrics need both measured values and uncertainties of those values. Travel monitoring smartphone apps can record people's travel behavior, but trip data quality is limited by sensor errors, user labeling rates and the accuracy of inference algorithms used for travel diary creation. We discuss the use of phone app recorded travel diary data to estimate energy consumption, and propose the use of propagation of variance to find error bars for such estimates. We define energy consumption for one trip as trip length times the energy intensity per distance unit of the travel mode used. We characterize trip length errors with relative error and inferred trip mode errors with confusion matrix columns. The resulting variances of each measurement are then propagated to the final calculated energy consumption. We tested our uncertainty methods on a dataset that used phone app data combined with prompted recall, consisting of 92,234 labeled trips for over 500,000 miles. Accounting for uncertainty using expected energy intensities and variance propagation gives a dataset-wide aggregate energy consumption percent error of about 9%, within one standard deviation from the truth. Future work could involve applying similar methods to other travel diary based metrics.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

Count Every Trip: Finding the Uncertainty in Energy Estimates Made from Inferred Travel Modes

To properly inform transport policy and infrastructure changes, transportation related metrics need both measured values and uncertainties of those values. Travel monitoring smartphone apps can record people's travel behavior, but trip data quality is limited by sensor errors, user labeling rates and the accuracy of inference algorithms used for travel diary creation. We discuss the use of phone app recorded travel diary data to estimate energy consumption, and propose the use of propagation of variance to find error bars for such estimates. We define energy consumption for one trip as trip length times the energy intensity per distance unit of the travel mode used. We characterize trip length errors with relative error and inferred trip mode errors with confusion matrix columns. The resulting variances of each measurement are then propagated to the final calculated energy consumption. We tested our uncertainty methods on a dataset that used phone app data combined with prompted recall, consisting of 92,234 labeled trips for over 500,000 miles. Accounting for uncertainty using expected energy intensities and variance propagation gives a dataset-wide aggregate energy consumption percent error of about 9%, within one standard deviation from the truth. Future work could involve applying similar methods to other travel diary based metrics.

ADVANCED PROPULSION SYSTEMS↗

Estimating Travel Energy Consumption Uncertainty Based on Inferred Travel Mode and Sensed Travel Length

To properly inform transport policy and infrastructure changes, transportation related metrics need both measured values and uncertainties of those values. Travel monitoring smartphone apps can record people's travel behavior, but trip data quality is limited by sensor errors, user labeling rates and the accuracy of inference algorithms used for travel diary creation. We discuss the use of phone app recorded travel diary data to estimate energy consumption, and propose the use of propagation of variance to find error bars for such estimates. We define energy consumption for one trip as trip length times the energy intensity per distance unit of the travel mode used. We characterize trip length errors with relative error and inferred trip mode errors with confusion matrix columns. The resulting variances of each measurement are then propagated to the final calculated energy consumption. We tested our uncertainty methods on a dataset that used phone app data combined with prompted recall, consisting of 92,234 labeled trips for over 500,000 miles. Accounting for uncertainty using expected energy intensities and variance propagation gives a dataset-wide aggregate energy consumption percent error of about 8%, within one standard deviation from the truth. Future work could involve applying similar methods to other travel diary based metrics.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Scaling open-weight large language models for hydropower regulatory information extraction: A systematic analysis

Information extraction from regulatory and technical documents using large language models (LLMs) involves practical trade-offs between extraction quality and computational cost. We evaluate eight open-weight LLMs spanning 0.6B–70B parameters on hydropower licensing documents and report deployment-oriented evidence under a unified extraction schema and evaluation protocol. Across the model set, we observe clear scale-dependent trends in both baseline extraction quality and the effectiveness of reflective reasoning (self-checking) under our fixed-prompt, no-augmentation setting. Mid-scale models often provide a favorable balance of accuracy and efficiency, whereas the smallest models show limited or inconsistent gains from the reasoning variants tested. Larger models achieve the highest overall F1 scores but incur substantially greater compute and infrastructure requirements. We further find that reliability failure modes can distort conventional metrics in this domain: in particular, high recall can coincide with systematic extraction errors when models fabricate values for fields that are absent from the source text, underscoring the importance of conservative null handling and evidence-grounded evaluation. Overall, our study provides a reproducible resource–performance comparison for open-weight LLM-based extraction in hydropower regulatory documentation and offers practical guidance for model selection under different deployment constraints.

Evaluation protocol↗

The 4DEnVar-based weakly coupled land data assimilation system for E3SM version 2

Abstract. A new weakly coupled land data assimilation (WCLDA) system based on the four-dimensional ensemble variational (4DEnVar) method is developed and applied to the fully coupled Energy Exascale Earth System Model version 2 (E3SMv2). The dimension-reduced projection four-dimensional variational (DRP-4DVar) method is employed to implement 4DVar using the ensemble technique instead of the adjoint technique. With an interest in providing initial conditions for decadal climate predictions, monthly mean anomalies of soil moisture and temperature from the Global Land Data Assimilation System (GLDAS) reanalysis from 1980 to 2016 are assimilated into the land component of E3SMv2 within the coupled modeling framework with a 1-month assimilation window. The coupled assimilation experiment is evaluated using multiple metrics, including the cost function, assimilation efficiency index, correlation, root-mean-square error (RMSE), and bias, and compared with a control simulation without land data assimilation. The WCLDA system yields improved simulation of soil moisture and temperature compared with the control simulation, with improvements found throughout the soil layers and in many regions of the global land. In terms of both soil moisture and temperature, the assimilation experiment outperforms the control simulation with reduced RMSE and higher temporal correlation in many regions, especially in South America, central Africa, Australia, and large parts of Eurasia. Furthermore, significant improvements are also found in reproducing the time evolution of the 2012 US Midwest drought, highlighting the crucial role of land surface in drought lifecycle. The WCLDA system is intended to be a foundational resource for research to investigate land-derived climate predictability.

58 GEOSCIENCES↗

Security of quantum position-verification limits Hamiltonian simulation via holography

We investigate the link between quantum position-verification (QPV) and holography established in [1] using holographic quantum error correcting codes as toy models. By inserting the “temporal” scaling of the AdS metric by hand via the bulk Hamiltonian interaction strength, we recover a toy model with consistent causality structure. This leads to an interesting implication between two topics in quantum information: if position-based verification is secure against attacks with small entanglement then there are new fundamental lower bounds for resources required for one Hamiltonian to simulate another.

AdS-CFT Correspondence↗

CarcSeq Measurement of Rat Mammary Cancer Driver Mutations and Relation to Spontaneous Mammary Neoplasia

Abstract The ability to deduce carcinogenic potential from subchronic, repeat dose rodent studies would constitute a major advance in chemical safety assessment and drug development. This study investigated an error-corrected NGS method (CarcSeq) for quantifying cancer driver mutations (CDMs) and deriving a metric of clonal expansion predictive of future neoplastic potential. CarcSeq was designed to interrogate subsets of amplicons encompassing hotspot CDMs applicable to a variety of cancers. Previously, normal human breast DNA was analyzed by CarcSeq and metrics based on mammary-specific CDMs were correlated with tissue donor age, a surrogate of breast cancer risk. Here we report development of parallel methodologies for rat. The utility of the rat CarcSeq method for predicting neoplastic potential was investigated by analyzing mammary tissue of 16-week-old untreated rats with known differences in spontaneous mammary neoplasia (Fischer 344, Wistar Han, and Sprague Dawley). Hundreds of mutants with mutant fractions ≥ 10−4 were quantified in each strain, most were recurrent mutations, and 42.5% of the nonsynonymous mutations have human homologs. Mutants in the mammary-specific target of the most tumor-sensitive strain (Sprague Dawley) showed the greatest nonsynonymous/synonymous mutation ratio, indicative of positive selection consistent with clonal expansion. For the mammary-specific target (Hras, Pik3ca, and Tp53 amplicons), median absolute deviation correlated with percentages of rats that develop spontaneous mammary neoplasia at 104 weeks (Pearson r = 1.0000, 1-tailed p = .0010). Therefore, this study produced evidence CarcSeq analysis of spontaneously occurring CDMs can be used to derive an early metric of clonal expansion relatable to long-term neoplastic outcome.

McKim, Karen L.↗

Characteristics of the IBEX Ribbon and Their Implications for a Source Region Outside the Heliopause

This paper presents a comprehensive exploration of the Interstellar Boundary Explorer energetic neutral atom (ENA) ribbon, focusing on its spatial and temporal variations over 14 yr. Methodological advancements, including a refined map modeling procedure and a new ribbon separation technique with appropriate error propagation, enable a detailed investigation of the ribbon’s features. Utilizing statistically robust metrics, this study reveals details of the ribbon across energy and time. Key findings include energy- and time-dependent variations in flux, angular radius, ribbon profile width, and higher moments. By applying these metrics, we reveal new complexity to the evolution of the ribbon over time, highlighting the nuanced relationship between it and the solar wind. Furthermore, the study examines for the first time the ribbon as it passes through the starboard/heliotail region (Lon EC 120°–180°), revealing properties distinct from other portions of the ribbon. The analysis uncovers an anticorrelation between ribbon width and flux, which provides quantitative support for a multisource ribbon created by a combination of solar wind neutrals that generate a spatiall narrow ribbon component and heliosheath neutrals giving rise to a broad component. Finally, differences in the temporal evolution of the ENA flux at different energies provide additional support that the location of the ribbon source region is beyond the heliopause.

79 ASTRONOMY AND ASTROPHYSICS↗