Engineering PapersSearch

SEARCH · Engineering Papers

Results for “validation data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The ARM Precipitation Best Estimate (PrecipBE) Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s Precipitation Best-Estimate (PrecipBE) Value-Added Product integrates multiple precipitation datastreams, accounting for data quality and instrument limitations, to deliver comprehensive per-precipitation event properties alongside ancillary ARM data set data. PrecipBE bundles all valid surface rainfall samples into artificial intelligence (AI)-ready tabular and time-series formats, reporting bundle means and uncertainty ranges. This per-event structure provides an insightful and easy-to-use resource for researchers analyzing precipitation characteristics.

54 ENVIRONMENTAL SCIENCES

Experimental Investigation of Vapor Formation in Liquid CO2 Flow Through a Converging-Diverging Nozzle

Carbon dioxide is an attractive working fluid for many cycles, including for pumped thermal energy storage (PTES). A challenge with some proposed sCO2 PTES cycles is the operation of sCO2 machinery outside the typical bounds of experience, with local phase change from the liquid state being particularly unknown. Presently, there is insufficient data in the literature regarding multiphase CO2 to adequately design a multiphase-tolerant turbine, so generation of foundational data is required. This experimental study investigates the flow characteristics of sub-sonic liquid CO2 undergoing expansion and phase change in a converging-diverging nozzle. The nozzle is instrumented to measure static pressure, unsteady pressure, temperature, and density. The static pressure transducers are located at 27 axial locations to accurately characterize the pressure profile in the nozzle. High-accuracy RTDs are located at the entrance and exit of the nozzle, and three dynamic pressure transducers are strategically located to capture any unsteady phenomena. During testing, values of mass flow and nozzle inlet pressure are swept to vary the pressure drop and fluid properties. The measured total pressure drop in the nozzle is compared to a homogenous model and the Lockhart-Martinelli correlation method, with the latter predicting loss quite closely. The resulting data set is valuable for validating multiphase numerical models in a simple geometry before implementation of these models in turbomachinery design.

25 ENERGY STORAGE

Data Compilation and Analysis from the Sirius-1 Experiment at TREAT for Transient Simulation Validation

This report assembles comprehensive data from the Sirius-1 experiment conducted by Idaho National Laboratory in collaboration with the National Aeronautics and Space Administration. The primary goal is to provide a robust data set that external users can utilize for the validation of computational methods for transient multiphysics simulations. By compiling all relevant data, including experiment design calculations, detailed engineering drawings for the experiment and data from reactor and fuel specimen measurement, this report is intended to serves as a reference for researchers and engineers working on the development and validation of computational models for transient nuclear behavior.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

WE-Validate: An Open-Source Framework For Wind Power Validation

Grid operators rely on historical weather time series at existing and planned wind power plants to make informed decisions when planning for a future power grid with very high penetration of renewable power. While synthetic wind power time series have been developed based on historical weather models, their validation with actual power production data remains complex due to variations in modeling practices and methodologies. This paper introduces the WE-Validate framework, originally designed for wind speed validation and now enhanced for wind power validation with a graphical user interface to support users with minimal programming experience. Validation of wind power with WE-Validate is based on robust metrics consisting of RMSE, centered RMSE, average bias, average percent bias, mean absolute error, mean absolute percent error, cross correlation, and calculation of ramping magnitude, rate, and duration. This paper showcases WE-Validate with validation of synthetically derived power for a wind plant in Washington state for one month in 2018. Validation of the synthetic power from two comparison data sets compared with observations shows both comparison series have strong correlation with observed across weekly and monthly aggregations while suffering from persistent negative bias. The suite of metrics within WE-Validate facilitates immediate insight into the utility of the comparison data sets through compression across multiple axes. This user-friendly, open-source tool can be extended beyond wind power, making it a valuable resource for system planners and operators in different domains.

Moncheur de Rieudotte, Malcolm P.

Dark Energy Survey Year 6 Results: Weak Lensing and Galaxy Clustering Cosmological Analysis Framework

We present the methodology for the weak lensing and galaxy clustering analyses of the Dark Energy Survey (DES) Year 6 data set. In this work, we design and validate the analysis pipeline for the cosmic shear, galaxy clustering plus galaxy$-$galaxy lensing ($2 \times 2$pt), and the joint analysis in the $3 \times 2$pt. Our framework accounts for key theoretical uncertainties, such as baryonic feedback and galaxy bias, incorporating both linear and non-linear models. We apply scale cuts in regimes where theoretical modeling becomes unreliable. The robustness of the pipeline is validated using mock data and simulations, confirming unbiased cosmological constraints and highlighting the importance of posterior projection effects in the validation process. As a result, we deliver robust and validated analysis pipelines for cosmic shear, $2 \times 2$pt, and $3 \times 2$pt in $Λ$CDM and $w$CDM scenarios, including a well-defined set of scales suitable for real data analysis, a robust prescription for theoretical systematics, and the theoretical covariance of the signal. This comprehensive methodology also lays the groundwork for future galaxy surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

Sanchez-Cid, D. [Zurich U.; Madrid, CIEMAT; Madrid

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]

Measuring the Conditional Luminosity and Stellar Mass Functions of Galaxies by Combining the Dark Energy Spectroscopic Instrument Legacy Imaging Surveys Data Release 9, Survey Validation 3, and Year 1 Data

In this investigation, we leverage the combination of the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Surveys Data Release 9, Survey Validation 3, and Year 1 data sets to estimate the conditional luminosity functions and conditional stellar mass functions (CLFs and CSMFs) of galaxies across various halo mass bins and redshift ranges. To support our analysis, we utilize a realistic DESI mock galaxy redshift survey (MGRS) generated from a high-resolution Jiutian simulation. An extended halo-based group finder is applied to both MGRS catalogs and DESI observation. By comparing the r- and z-band luminosity functions (LFs) and stellar mass functions (SMFs) derived using both photometric and spectroscopic data, we quantified the impact of photometric redshift (photo-z) errors on the galaxy LFs and SMFs, especially in the low-redshift bin at the low-luminosity/mass end. By conducting prior evaluations of the group finder using MGRS, we successfully obtain a set of CLF and CSMF measurements from observational data. We find that at low redshift, the faint-end slopes of CLFs and CSMFs below ~10 9 h –2 L ⊙ (or h –2 M ⊙ ) evince a compelling concordance with the subhalo mass functions. After correcting the cosmic variance effect of our local Universe following Chen et al., the faint-end slopes of the LFs/SMFs turn out to also be in good agreement with the slope of the halo mass function.

79 ASTRONOMY AND ASTROPHYSICS

Identification of Distorted Gamma-Ray Signature Patterns Using Digital Filtering and Auto-Associative Memory Implemented with a Hopfield Neural Network

The detection and identification of radioactive sources in search applications involve analyzing passive gamma-ray emissions from high-level radioactive materials. This process uses a mobile detector-spectrometer in a complex field test environment. Recently, the use of artificial intelligence for gamma-ray spectrum analysis has shown promising results. However, challenges persist in identifying isotopic signatures from spectral measurements that may be distorted due to source shielding, random variations in natural radioactive background, or insufficient measurement time to obtain clear spectral lines. Here, this paper presents a novel intelligent signature recognition method that combines digital filtering techniques with an artificial Hopfield Neural Network (HNN). The HNN leverages auto-associative memory to store training sample patterns and match them with incoming gamma spectra from distorted sources. It restores the testing sources’ measurements by finding the closest matching signature patterns in the spectral library. Before HNN recognition, the measured spectrum undergoes preprocessing with a digital image filter to reduce fluctuations. Performance of the proposed method is evaluated using a set of gamma-ray spectra measured with a sodium iodide detector. The data collected include measurements from six pure samples: 241 Am, 60 Co, 137 Cs, 192 Ir, 239 Pu, and 235 U, which are used for training and validation (i.e. six cases). Additionally, the data set contains 24 distorted synthesized sources with various fluctuating backgrounds. Test results demonstrate the potential of the proposed method to accurately recognize the correct isotope with high precision, achieving an accuracy rate exceeding 85%. Furthermore, the proposed method exhibits superior performance compared to the conventional multiple regression fitting and simple feedforward neural network methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

A review of NEST models for liquid xenon and an exhaustive comparison with other approaches

This paper discusses the microphysical simulation of interactions in liquid xenon, the active detector medium in many leading rare-event searches for new physics, and describes experimental observables useful for understanding detector performance. The scintillation and ionization yield distributions for signal and background are presented using the Noble Element Simulation Technique (NEST), a toolkit based on experimental data and simple empirical formulas, which mimic previous microphysics modeling but are guided by data. The NEST models for light and charge production as a function of the particle type, energy, and electric field are reviewed, along with models for energy resolution and final pulse areas. NEST is compared with other models or sets of models and validated against real data, with several specific examples drawn from XENON, ZEPLIN, LUX, LZ, PandaX, and table-top experiments used for calibrations.

WIMPs

Predicting RNA structure and dynamics with deep learning and solution scattering

Advanced deep learning and statistical methods can predict structural models for RNA molecules. However, RNAs are flexible, and it remains difficult to describe their macromolecular conformations in solutions where varying conditions can induce conformational changes. Small-angle x-ray scattering (SAXS) in solution is an efficient technique to validate structural predictions by comparing the experimental SAXS profile with those calculated from predicted structures. There are two main challenges in comparing SAXS profiles to RNA structures: the absence of cations essential for stability and charge neutralization in predicted structures and the inadequacy of a single structure to represent RNA’s conformational plasticity. We introduce a solution conformation predictor for RNA (SCOPER) to address these challenges. This pipeline integrates kinematics-based conformational sampling with the innovative deep learning model, IonNet, designed for predicting Mg 2+ ion binding sites. Validated through benchmarking against 14 experimental data sets, SCOPER significantly improved the quality of SAXS profile fits by including Mg 2+ ions and sampling of conformational plasticity. We observe that an increased content of monovalent and bivalent ions leads to decreased RNA plasticity. Therefore, carefully adjusting the plasticity and ion density is crucial to avoid overfitting experimental SAXS data. SCOPER is an efficient tool for accurately validating the solution state of RNAs given an initial, sufficiently accurate structure and provides the corrected atomistic model, including ions.

59 BASIC BIOLOGICAL SCIENCES

MARIAH PCAP data for Validation Demonstration

This dataset holds simulated PCAP (packet capture) data from the SCEPTRE validation demonstration model as a set of pairwise communications between devices via specific protocols. All connections should be assumed to be symmetric, as this data is an aggregation of the true PCAP. A mapping is also provided associating each IP address with its true device type.

cyber-physical system

Beyond Optimization: Exploring Novelty Discovery in Autonomous Experiments

Autonomous experiments (AEs) are transforming how scientific research is conducted by integrating artificial intelligence with automated experimental platforms. Current AEs primarily focus on the optimization of a predefined target; while accelerating this goal, such an approach limits the discovery of unexpected or unknown physical phenomena. Here, we introduce a novel framework, INS 2 ANE (Integrated Novelty Score−Strategic Autonomous Non-Smooth Exploration), to enhance the discovery of novel phenomena in autonomous microscopy experimentation. Our method integrates two key components: (1) a novelty scoring system that evaluates the uniqueness of experimental results and (2) a strategic sampling mechanism that promotes exploration of under-sampled regions even if they appear less promising by conventional criteria. We validate this approach on a preacquired data set with a known ground truth comprising of image−spectral pairs. We further implement the process on autonomous scanning probe microscopy experiments. INS 2 ANE significantly increases the diversity of explored phenomena in comparison to conventional optimization routines, enhancing the likelihood of discovering previously unobserved phenomena. These results demonstrate the potential for autonomous microscopy experiments to enhance the scientific discovery by navigating complex experimental spaces to uncover novel phenomena.

Materials

Validation of the DESI 2024 Lyα forest BAO analysis using synthetic datasets

The first year of data from the Dark Energy Spectroscopic Instrument (DESI) contains the largest set of Lyman-α (Lyα) forest spectra ever observed. This data, collected in the DESI Data Release 1 (DR1) sample, has been used to measure the Baryon Acoustic Oscillation (BAO) feature at redshift z = 2.33. In this work, we use a set of 150 synthetic realizations of DESI DR1 to validate the DESI 2024 Lyα forest BAO measurement presented in [1]. The synthetic data sets are based on Gaussian random fields using the log-normal approximation. We produce realistic synthetic DESI spectra that include all major contaminants affecting the Lyα forest. The synthetic data sets span a redshift range 1.8 < z < 3.8, and are analyzed using the same framework and pipeline used for the DESI 2024 Lyα forest BAO measurement. To measure BAO, we use both the Lyα auto-correlation and its cross-correlation with quasar positions. We use the mean of correlation functions from the set of DESI DR1 realizations to show that our model is able to recover unbiased measurements of the BAO position. We also fit each mock individually and study the population of BAO fits in order to validate BAO uncertainties and test our method for estimating the covariance matrix of the Lyα forest correlation functions. Finally, we discuss the implications of our results and identify the needs for the next generation of Lyα forest synthetic data sets, with the top priority being to simulate the effect of BAO broadening due to non-linear evolution.

79 ASTRONOMY AND ASTROPHYSICS

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

Quantifying Groundwater Response and Uncertainty in Beaver‐Influenced Mountainous Floodplains Using Machine Learning‐Based Model Calibration

Abstract Beavers ( Castor canadensis ) alter river corridor hydrology by creating ponds and inundating floodplains, and thereby improving surface water storage. However, the impact of inundation on groundwater, particularly in mountainous alluvial floodplains with permeable gravel/cobble layers overlain by a soil layer, remains uncertain. Numerical modeling across various floodplain structures considers topographic and sediment complexity and multidirectional flow, linking inundation to groundwater response. This study develops a model‐data integration workflow to address uncertainty in groundwater response to beaver‐induced inundations in a mountainous alluvial floodplain in the Upper Colorado River Basin. Uncertain factors include seasonal hydrologic dynamics, hydraulic conductivities, floodplain structures, and meteorological forcings. We employed an ensemble of groundwater models, based on geophysical and hydrologic data, with machine learning‐based calibration using a neural density estimator. This allowed us to quantify the vertical flux from the soil layer to the permeable gravel bed, the down‐valley underflow within the gravel bed, and their ratios. Results show a significant increase in the vertical flux relative to down‐valley underflow, from 2 during dry pond periods to 20 during wet periods, serving as an analogy for conditions without and with beaver ponds. The study highlights the influence of floodplain structure on groundwater storage, water balance, and water quality impacted by beaver ponds. A thick gravel bed layer, with a large down‐valley underflow, minimizes the effect of beaver‐induced inundation on water quality. We emphasize the need for field‐scale measurements of floodplain structure and improved characterization of evapotranspiration changes to reduce uncertainty in groundwater response. Plain Language Summary Beavers change the flow of water in river corridors by creating ponds, expanding wetlands, and flooding floodplains. This increases surface water area, promotes plant growth, and enhances biodiversity. However, the impact of this flooding on groundwater flow is not well understood, especially in mountainous areas with gravel layers where water moves easily beneath soil. In this study, we used numerical modeling to investigate how beaver ponds influence groundwater in a mountainous floodplain of the Upper Colorado River Basin. We adapted a machine learning method to validate our numerical models using multiple field data sets. Our findings show that beaver ponds significantly increase vertical water flow from the soil to the gravel during wet periods, compared to when the ponds are fully drained. The study also highlights the importance of floodplain structure in controlling both water flow in gravel layers along the river direction and vertical flow from the soil to the gravel with the presence of beavers. To reduce uncertainty in groundwater response, we emphasize the need for more field‐scale measurements of floodplain structure, hydraulic properties, and evapotranspiration changes. Key Points Floodplain structures and hydraulic conductivities are important for groundwater response with beaver ponds in mountainous floodplains Large down‐valley underflow in permeability‐stratified floodplains reduces beaver‐induced impacts on groundwater storage and water quality Machine learning‐based model calibration methods are effective for estimating posterior distributions of groundwater model parameters

Wang, Lijing

Avoidance of disruptions on KSTAR due to vertical displacement events via novel real-time stability assessment

Disruption avoidance via the DECAF approach has been achieved on KSTAR using a novel real-time vertical stability assessment and a multiactuator feedback control strategy. The development of disruption avoidance strategies with reactor-relevant reliability is an urgent activity, enabling future fusion power plants. The stability metric employed is based on a new formulation of a vertical force gradient balance metric evaluated across the poloidal cross section of the plasma, with parameters tuned using historical data. Evaluation of this metric on a validation set of 400 recent KSTAR shots indicates >82% of Vertical displacement events can be avoided via feedback control. Essential to its calculation is the two-dimensional toroidal current density distribution in the plasma. Measurement of this profile faster than fully-converged equilibrium reconstructions can deliver is found to improve forecaster performance and is achieved with a surrogate model that takes as input magnetic diagnostic measurements and outputs the current profile on a basis comprising the top principal components of historical current profiles (from past equilibrium reconstructions). This method solves the non-uniqueness problem typically faced when reconstructing current profiles directly from diagnostics, while improving computational time and accuracy. On average, profiles produced by this model reach coefficients of determination of >0.99 with respect to those from equilibrium reconstructions. The avoidance actuators employed include poloidal field coils and an electron cyclotron current drive system. The multiactuator approach, as shown in this first demonstration, allows disruption avoidance while minimizing impact to operational performance. This ability, along with its flexibility and speed, makes this new approach an attractive option for avoiding these types of disruptions in reactors.

Tobin, Matthew [Columbia Univ., New York, NY (Unit

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION