Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Experimental data for damage mechanics simulation challenge

While there are many computational approaches for simulating damage in rock and other materials, few have been ground truth tested with either known experimental data or with blind data sets. Here, in this work, we present a bench-mark laboratory data set for a damage mechanics challenge to compare computational approaches on damage evolution in brittle-ductile materials. The samples were fabricated through additive manufacturing to produce repeatable specimens designed to fail in controlled ways. The failure was induced in the samples using a 3-point bending test to produce different Modes such as Mode I and mixed Modes including I-II, I-III and I-II-III Modes to generate a calibration data set and a blind challenge data set. Data collected included spatial and temporal measurements from traditional digital load–displacement sensors, 2D digital image correlation measurement to map surface deformations, 3D X-ray microscopy to ground-truth the crack-failure geometry, and laser profilometry to capture surface roughness. The data sets are available, on a data repository, to the community to advance computational models to improve our ability to predict damage in brittle-ductile materials.

3-point bending

Quality Ranking of Unary Chloride Salt Property Data Included in MSTDB-TP

Molten salt reactor developers rely on thermal property data to design, license and operate the reactors. The Molten Salt Thermal Database-Thermophysical Properties (MSTDB-TP) was established under the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program and is managed by Oak Ridge National Laboratory to serve as a single source of thermophysical property values measured for a wide variety of molten salt systems for use by researchers, molten salt reactor developers, and regulators. These properties include density, viscosity and thermal diffusivity and conductivity. Published measurements of molten salt properties are lacking for many salts of interest and the data that are available are often inconsistent. This creates a challenge for MSR developers when determining which property values to use when designing their reactors. It is the purpose of this work to apply a consistent ranking system to all data entries that indicates the quality of property values listed in the database. These rankings will be the technical basis for down-selections by the database developers and alert users about the quality of the available property values. MSTDB-TP collects all available property data and indicates preferred data sets or correlations. However, all available data sets are included in the database. Quality assessments and rankings are being applied to data in MSTDB-TP to provide an indication of the quality of each data set independent of consistency with other data. Previous reports detailed the ranking system that was followed and assessments of unary fluoride data sets. Documentation of the quality of data in MSTDB-TP was continued by reviewing and assessing all available sources of density, viscosity and thermal diffusivity or conductivity values for unary chloride salts in MSTDB-TP V3.0 using the same criteria.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5

BENEFIT with Northeastern University: HVAC Hardware-in-the-Loop Experimental Testing of a Heat Pump and Air Conditioner

This dataset includes HVAC Hardware-in-the-Loop (HIL) experimental results for a single stage, SEER 16, HSPF 9.5, 3-ton single-speed air source heat pump with 15 kW of backup auxiliary heating tested in both cooling and heating mode, and a two stage, SEER 21, 2-ton central air conditioner tested in cooling mode for a set of outdoor temperatures and indoor setpoint temperatures. In addition to these tests, experimental tests focused on the operation of auxiliary heating for the heat pump for winter condition were also conducted. The laboratory experiments for transient testing of the heat pump and air conditioner were conducted using the two HIL systems in the Systems Performance Laboratory (SPL) at NREL’s Energy Systems Integration Facility (ESIF). Further information on laboratory design and capabilities of the SPL along with the architecture of HVAC HIL system can be found in: Sparn, B. F. 2018. Laboratory Resources and Techniques to Evaluate Smart Home Technology (No. NREL/CP-5500-71696). National Renewable Energy Laboratory (NREL), Golden, CO (United States). https://www.nrel.gov/docs/fy18osti/71696.pdf and the experimental setup and validation of HVAC HIL platform can be found in: Ramaraj, S. and Sparn, B. 2022. Validation of HVAC Hardware-In-the-Loop Simulation for Advanced Control Strategies in Smart Homes (No. NREL/CP-5500-82562). National Renewable Energy Lab (NREL), Golden, CO (United States). https://www.nrel.gov/docs/fy22osti/82562.pdf. These experimental results can be used to validate how we currently model the cycling behavior of heat pumps and air conditioners. Additionally, many demand response programs implement heat pump and air conditioner control by changing the thermostat set point – these data may also be used to verify our models for heat pump and air conditioner demand response control are implemented correctly. The Test_Matrix file describes all the indoor and outdoor test conditions for heat pump and air conditioner and the file names of data sets include information about the test conditions. A wide range of outdoor air temperatures were chosen to accommodate summer and winter conditions. In addition to operating the HVAC equipment with different outdoor temperatures, we also operate the system with different indoor temperature set points to represent different grid signals or different operating conditions. For cooling conditions, the baseline set point is 72°F. To represent Load Up signals, the setpoint is changed to 68°F. The Load Shed set point is 76°F. For heating conditions, the baseline set point was assumed to be 68°F. The Load add set point is 72°F and the Load shed set point is 64°F. The starting indoor temperature for cooling conditions was set ~2°F above the indoor setpoint temperature so that the equipment turned on quickly. Similarly, the initial indoor temperature was set ~2°F lower than setpoint for heating mode tests to ensure that heating began quickly. The return air temperature was assumed to be equal to the indoor setpoint temperature in all cases. The experimental data are sampled at 1-second intervals. The data from ecobee thermostat at 5-minute interval are resampled and added to the corresponding file. The content of each data set is as follows: • T_Return (C): Measured return air temperature [C] • T_Return_SP (C): Return air temperature setpoint from E+ model, sent to HIL [C] • T_Supply (C): Measured supply air temperature at evaporator outlet [C] • T_Outdoor (C): Measured outdoor air temperature [C] • T_Outdoor_SP (C): Outdoor air temperature setpoint from weather file, sent to HIL [C] • T_Indoor (C): Measured indoor air temperature [C] • T_Indoor_SP (C): Indoor air temperature setpoint from E+ model, sent to HIL [C] • Outdoor Unit Power (W): Measured power of the outdoor unit [W] • Indoor Unit Power (W): Measured power of the indoor unit [W] • Evaporator Airflow Rate (CFM): Measured evaporator or indoor unit airflow rate sent to E+ model [CFM] • Cooling/Heating Capacity (kW): Calculated cooling/heating capacity sent to E+ model [kW] • T_SP_Thermostat (C): Thermostat cooling/heating setpoint temperature [C] • T_Indoor_Thermostat (C): Thermostat indoor air temperature [C]

24 POWER TRANSMISSION AND DISTRIBUTION

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

Measuring Neutron Polarisation in Deuteron Photo-disintegration with the CLAS Start Counter [Thesis]

Deuteron photo-disintegration (γd → γp) is a reaction that represents the simplest case in which nuclear and hadron physics models can be tested. Despite this, associated polarization analyses are limited in terms of angular coverage and energy ranges, especially in observables related to the recoil neutron. This is largely due to a lack in dedicated polarimetry equipment, and represents a roadblock in global progress to understand high-energy phenomena such as hexaquarks, and quark-gluon degrees of freedom. To address this problem, this PhD thesis pioneers a new methodology for the parasitic measurement of nucleon polarization using kinematic reconstruction of (spin-dependent) nucleon-nucleus scattering of reaction products, prior to their detection in large acceptance particle detector apparatus. Following this novel approach, which requires no dedicated polarimeter, a determination of the double polarization observable, $C^n_{x'}$, from deuteron photo-disintegration is presented, using Jefferson Lab’s CLAS detector. The analysis utilizes the (n,p) charge exchange reaction in CLAS’s "start counter" (plastic scintillator) to determine the final state neutron polarizations. The results present the first ever data for this observable above 0.7 GeV (photon beam energy) and significantly extend the angular range of the world data set. This new data is largely statistically consistent with the previous measurement of $C^n_{x'}$ by Bashkanov et al . in the overlapping energy range of 0.4-0.7 GeV. It is planned for the statistical accuracy of the presented result to be increased by the inclusion of additional data. The analysis herein serves as a key proof of concept for future applications, including a recommended similar analysis to be implemented with data from the more modern CLAS12 detector. This paves the way for a plethora of additional analyses using existing data sets that would provide crucial new constraints for hadron and nuclear physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Satellite Imagery of PV Site Storm Damage

"This repository contains multiple data sets focused on visible damage to photovoltaic (PV) installations following extreme weather events such as hailstorms and hurricanes. Data sets are split into two categories: the first category, the ‘manually labeled’ data, was compiled by researchers manually, and contains manually identified PV sites exposed to storms. The second data set, the ‘aggregated’ data, is a compilation of the manually labeled PV sites and deep learning-identified PV sites. The hail damage data set focuses on post-storm PV damage following a September 24, 2023 hailstorm in Austin, TX, which caused over $600 million in damages in the Austin metro area. The hurricane damage data set focuses on post-storm PV damage following Hurricanes Irma and Maria in Puerto Rico and the US Virgin Islands. Hurricanes Irma and Maria were back-to-back category 5 hurricanes, which pummeled the Caribbean and southeastern United States in September 2017, causing an estimated $115.2 billion in damages."

14 SOLAR ENERGY

Long‐Term Large‐Scale Atmospheric Forcing Data From Three‐Dimensional Constrained Variational Analysis for the ARM SGP Site

Here, this study presents a long‐term three‐dimensional large‐scale forcing data set (VARANAL3D) derived from the three‐dimensional constrained variational analysis (3DCVA) method at the Atmospheric Radiation Measurement (ARM) program Southern Great Plains (SGP) site from 2004 to 2018. Building on the same input data sets as the conventional continuous forcing data set (VARANAL), VARANAL3D maintains overall consistency in domain‐averaged fields while introducing spatial variability, offering critical insights into the influence of mesoscale synoptic systems on cloud‐related processes. Evaluations are conducted across four cloud and precipitation regimes: Clear‐sky, Shallow‐clouds, Afternoon‐precipitation, and Nocturnal‐precipitation, presenting high consistency of the domain‐mean forcing data sets while emphasizing the role of subdomain forcing variability particularly in precipitating regimes. Single column model (SCM) simulations demonstrate that subdomain VARANAL3D forcing improves cloud and precipitation representation, with the ensemble outperforming domain‐mean forcing in three cloudy and precipitating regimes. Overall, these results highlight VARANAL3D's value for investigating the impacts of spatial variability of large‐scale forcing on atmospheric processes. The VARANAL3D data set provides new opportunities for evaluating model physics, advancing the development of scale‐aware parameterizations and deepening our understanding of cloud and precipitation dynamics.

Environmental sciences

Historical and Future Global Irrigation Energy Consumption by Fuel and Region

Irrigation energy use is a significant component of agricultural production costs, contributing directly to the energy and emissions intensity of crop production and ultimately to food prices. Understanding the existing structure of irrigation energy consumption help achieve food-energy-water security and environmental goals. We present a comprehensive global data set detailing country-level irrigation energy consumption, emphasizing the comparative use of electric, diesel, and emerging solar pumps. To our knowledge, no such data set exists. We draw from a literature review to develop a logistic transformed regression model to estimate the shares of fuel sources for irrigation across countries over historical years to construct a global data set of country-level irrigation energy consumption by multiple fuel sources. Additionally, we compare our estimates of irrigation energy use with agricultural energy use as reported by the International Energy Agency and other external sources. We then use this data to project future irrigation energy use with the Global Change Analysis Model, which is a multisector dynamics model, to showcase the usage of this data set. Projections under the reference scenario show a global shift in fuel types for irrigation pumping, while patterns vary across regions, with India and Pakistan leading in solar-powered irrigation growth and countries like the USA and China continuing to rely primarily on grid electricity. This data set provides a resource to understand the role of irrigation fuel choices within the broader energy sector, as well as the connected agricultural, land use, and water sectors under alternative future scenarios, enabling informed decision making toward efficient agricultural practices.

Global Change Analysis Model (GCAM)

SPRUCE Ground Observations of Phenology in Experimental Plots, 2024

This data set consists of one comma separated (*.csv) file containing phenological transition dates, as derived from direct observations of vegetative and reproductive phenology recorded by a human observer, from the SPRUCE experiment during 2024 (2025-03-06 to 2025-11-21), the ninth full year of whole-ecosystem warming (Hanson et al. 2017). Both spring and autumn phenological events are included. Since April 2016, human observers have been directly tracking the phenology of both woody and herbaceous species on a weekly schedule within the SPRUCE experimental chambers, these data are reported in annual ground observations data sets (see Related Data Sets). The observed date reported here is the first survey date in 2024 on which an event/phenophase was definitively observed. This data set also contains a companion file in HTML (*.html) containing figures showing the relationship between the day of year and temperature treatment for different phenological phases by species for 2024.

54 ENVIRONMENTAL SCIENCES

Measurements of the thermal and ionization state of the intergalactic medium during the cosmic afternoon

We perform the first measurement of the thermal and ionization state of the intergalactic medium (IGM) across 0.9 < z < 1.5 using 301 Ly α absorption lines fitted from 12 archival Hubble Space Telescope Space Telescope Imaging Spectrograph quasar spectra. We employ the machine-learning-based inference method that uses joint Doppler parameter–column density (⁠b-N HI ⁠) distributions obtained from Ly α forest decomposition. Our results show that the Γ HI photoionization rates, ⁠, agree with recent ultraviolet background synthesis models, with log(Γ HI /s -1 ) = $-11.79^{+0.18}_{-0.15}$, $-11.98^{+0.09}_{-0.09}$⁠, and $-12.32^{+0.10}_{-0.12}$⁠, at z = 1.4, 1.2, and 1, respectively. We obtain the IGM temperature at the mean density, T 0 ⁠, and the adiabatic index, γ⁠, as [log(T 0 /K), γ] = $[4.13^{+0.12}_{-0.10}, 1.34^{+0.10}_{-0.15}]$, $[3.79^{+0.11}_{-0.11}, 1.70^{+0.09}_{-0.09}]$, and $[4.12^{+0.15}_{-0.25}, 1.34^{+0.21}_{-0.26}]$ at z = 1.4⁠, 1.2, and 1. Our measurements of T 0 at z = 1.4 and 1.2 are consistent with the trend predicted from previous z < 3 temperature measurements and theoretical expectations, where the IGM cools down after $He\tiny{II}$ reionization in the absence of any non-standard heating. However, our T 0 measurement at z = 1 unexpectedly high IGM temperature. Given the relatively large uncertainty in these measurements, where σ T$_0$ ~ 5000 K, mostly emanating from the limited size of our data set, we cannot conclude whether the IGM cools down as expected. Lastly, we generate mock data sets to test the constraining power of future measurement with larger data sets. The results demonstrate that, with redshift path-length Δz ~ 2 for each redshift bin, three times the current data set, we can constrain the T 0 of IGM within 1500 K, which would be sufficient to constrain the IGM thermal history at z < 1.5 conclusively.

79 ASTRONOMY AND ASTROPHYSICS

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY

Digital image correlation and infrared thermography data for seven unique geometries of 304L stainless steel

Material Testing 2.0 (MT2.0) is a paradigm that advocates for the use of rich, full-field data, such as from digital image correlation and infrared thermography, for material identification. By employing heterogeneous, multi-axial data in conjunction with sophisticated inverse calibration techniques such as finite element model updating and the virtual fields method, MT2.0 aims to reduce the number of specimens needed for material identification and to increase confidence in the calibration results. To support continued development, improvement, and validation of such inverse methods—specifically for rate-dependent, temperature-dependent, and anisotropic metal plasticity models—we provide here a thorough experimental data set for 304L stainless steel sheet metal. The data set includes full-field displacement, strain, and temperature data for seven unique specimen geometries tested at different strain rates and in different material orientations. Commensurate extensometer strain data from tensile dog bones is provided as well for comparison. We believe this complete data set will be a valuable contribution to the experimental and computational mechanics communities, supporting continued advances in material identification methods.

36 MATERIALS SCIENCE