Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES

A Cloud-Tracking Data Set for the CSAPR2 Adaptive Scanning during TRACER

The U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) User Facility (Mather and Voyles 2013) deployed the first ARM Mobile Facility (AMF1; Miller et al. 2016) near LaPorte, Texas to support the Tracking Aerosol Convection Interactions Experiment (TRACER) (Jensen et al. 2025) near Houston, Texas. From October 2021 to September 2022, AMF1 was deployed to 29.67° N, 95.06° W near LaPorte, Texas and the 2nd Generation C-band Scanning ARM Precipitation Radar (CSAPR2) was deployed to a supplementary site at 29.53° N, 95.28° W (Figure 1). During an intensive operational period (IOP) from 1 June to 30 September 2022, the CSAPR2 sampled precipitation echoes in an adaptive scanning mode following the Multisensor Agile Adaptive Scanning (MAAS) framework (Kollias et al. 2020). MAAS helped optimize the CSAPR2 scan strategy to perform frequent plan position indicator (PPI) and range height indicator (RHI) scans (Lamer et al. 2023). Details of the CSAPR2 scanning, data processing, and calibration procedures used by the principal investigator (PI), and the PI data files are described by Oue et al. (2023). Details of the CSAPR2 operational performance, ARM data processing and correction procedures, and data quality masks are described by Feng et al. (2024a).

54 ENVIRONMENTAL SCIENCES

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie

Advanced Materials & Manufacturing Technology (AMMT): Development of Additive Manufacturing Agnostic Process Parameter Procedure, 316H Stainless Steel Readiness Level Data Sets, and Machine Maintenance Plan

The University of California, Davis is involved in a project to deploy and enhance an artificial intelligence (AI) system for predicting and preventing plasma disruptions on the DIII D tokamak, under the funding from Department of Energy DE-SC0023500 (title: AI/Deep Learning FRNN Software for Prediction & Real-Time Control of DIII-D Plasma Control System (PCS)). The overarching goal is to demonstrate that real-time, AI-guided intervention can proactively modify the plasma state to avoid or mitigate disruptions—a critical challenge for the future of fusion energy.

36 MATERIALS SCIENCE

The ARM Precipitation Best Estimate (PrecipBE) Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s Precipitation Best-Estimate (PrecipBE) Value-Added Product integrates multiple precipitation datastreams, accounting for data quality and instrument limitations, to deliver comprehensive per-precipitation event properties alongside ancillary ARM data set data. PrecipBE bundles all valid surface rainfall samples into artificial intelligence (AI)-ready tabular and time-series formats, reporting bundle means and uncertainty ranges. This per-event structure provides an insightful and easy-to-use resource for researchers analyzing precipitation characteristics.

54 ENVIRONMENTAL SCIENCES

SPRUCE Vegetation Phenology in Experimental Plots from PhenoCam Imagery, 2015-2024

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2025 (2015-08-24 to 2025-03-31), with start- and end-of-season phenological transition dates derived through the end of autumn 2024. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step. • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e., vegetation type). • Contains one file in *.csv format. (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure. • Contains two files in *.csv format, one for snow on trees and one for snow on ground. This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type. • One additional file in HTML format with the transition dates plotted for each vegetation type, by year. (2) R files for processing PhenoCam files and flags. • Contains five files in R file(*.R) format and the components of the phenocamr package (Version 1.1.4) used for calculating transition dates for 2015-2024. These are contained in a compressed (*.zip) file. User Note: All imagery is posted in near-real time to the PhenoCam Project web page (https://phenocam.nau.edu), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. This data set is based on the complete camera record from SPRUCE and supersedes all previously released PhenoCam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

54 ENVIRONMENTAL SCIENCES

Cambium 2024 Scenario Descriptions and Documentation

The National Renewable Energy Laboratory's (NREL's) Cambium data sets are annually released sets of simulated hourly data for a range of modeled futures of the U.S. electric sector with metrics designed to be useful for long-term decision- making. The 2024 Cambium data set is the fifth annual release. The data sets are a companion product to NREL's Standard Scenarios, which are likewise released annually and are a set of projections of how the U.S. electric sector could evolve across a suite of different potential futures, but covering more scenarios with less temporal granularity. Information about Cambium and related publications can be found at https://www.nrel.gov/analysis/cambium.html, and the Cambium data sets can be viewed and downloaded at https://scenarioviewer.nrel.gov/. In this documentation, we describe Cambium 2024's scenarios, define the metrics, and document the Cambium-specific methods for calculating those metrics.

24 POWER TRANSMISSION AND DISTRIBUTION

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Oak Ridge National Laboratory EAGLE-I TM : Modeling Electric Utility County Customers for Situational Awareness

During natural hazard events (hurricanes, wildfires, earthquakes, etc.) and recent man-made events (e.g., cyber attacks), the exchange of near real-time, spatially refined data within the response community is critical. The EAGLE-I$^{TM}$ platform is one tool that facilitates this data for decision makers within the energy sector. While much information can be collected and integrated into the system directly, other pertinent data must be augmented by other derived data products to enhance the information and allow for a consistent evaluation of on-the-ground conditions. One such data set that requires the addition of other derived data is the electric utility customer outage data that is aggregated to the county level within the EAGLE-I application. Without a county customer data set, outages can only be compared on total counts, which gives greater importance to higher population outages. Including an electric utility customer data set at the county level allows for these outage counts to be converted to percent outages and brings a consistent classification of outages and equal importance to all outages. To achieve this, several available data sets were combined and spatial disaggregation techniques were employed to model customer estimates at the county scale. This paper presents the approach to produce this data for the United States and lessons learned from working with these disparate data sets. Data validation is provided, where possible, and limitations of the model and possible improvements are discussed.

24 POWER TRANSMISSION AND DISTRIBUTION

Experimental data for damage mechanics simulation challenge

While there are many computational approaches for simulating damage in rock and other materials, few have been ground truth tested with either known experimental data or with blind data sets. Here, in this work, we present a bench-mark laboratory data set for a damage mechanics challenge to compare computational approaches on damage evolution in brittle-ductile materials. The samples were fabricated through additive manufacturing to produce repeatable specimens designed to fail in controlled ways. The failure was induced in the samples using a 3-point bending test to produce different Modes such as Mode I and mixed Modes including I-II, I-III and I-II-III Modes to generate a calibration data set and a blind challenge data set. Data collected included spatial and temporal measurements from traditional digital load–displacement sensors, 2D digital image correlation measurement to map surface deformations, 3D X-ray microscopy to ground-truth the crack-failure geometry, and laser profilometry to capture surface roughness. The data sets are available, on a data repository, to the community to advance computational models to improve our ability to predict damage in brittle-ductile materials.

3-point bending

Quality Ranking of Unary Chloride Salt Property Data Included in MSTDB-TP

Molten salt reactor developers rely on thermal property data to design, license and operate the reactors. The Molten Salt Thermal Database-Thermophysical Properties (MSTDB-TP) was established under the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program and is managed by Oak Ridge National Laboratory to serve as a single source of thermophysical property values measured for a wide variety of molten salt systems for use by researchers, molten salt reactor developers, and regulators. These properties include density, viscosity and thermal diffusivity and conductivity. Published measurements of molten salt properties are lacking for many salts of interest and the data that are available are often inconsistent. This creates a challenge for MSR developers when determining which property values to use when designing their reactors. It is the purpose of this work to apply a consistent ranking system to all data entries that indicates the quality of property values listed in the database. These rankings will be the technical basis for down-selections by the database developers and alert users about the quality of the available property values. MSTDB-TP collects all available property data and indicates preferred data sets or correlations. However, all available data sets are included in the database. Quality assessments and rankings are being applied to data in MSTDB-TP to provide an indication of the quality of each data set independent of consistency with other data. Previous reports detailed the ranking system that was followed and assessments of unary fluoride data sets. Documentation of the quality of data in MSTDB-TP was continued by reviewing and assessing all available sources of density, viscosity and thermal diffusivity or conductivity values for unary chloride salts in MSTDB-TP V3.0 using the same criteria.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5

BENEFIT with Northeastern University: HVAC Hardware-in-the-Loop Experimental Testing of a Heat Pump and Air Conditioner

This dataset includes HVAC Hardware-in-the-Loop (HIL) experimental results for a single stage, SEER 16, HSPF 9.5, 3-ton single-speed air source heat pump with 15 kW of backup auxiliary heating tested in both cooling and heating mode, and a two stage, SEER 21, 2-ton central air conditioner tested in cooling mode for a set of outdoor temperatures and indoor setpoint temperatures. In addition to these tests, experimental tests focused on the operation of auxiliary heating for the heat pump for winter condition were also conducted. The laboratory experiments for transient testing of the heat pump and air conditioner were conducted using the two HIL systems in the Systems Performance Laboratory (SPL) at NREL’s Energy Systems Integration Facility (ESIF). Further information on laboratory design and capabilities of the SPL along with the architecture of HVAC HIL system can be found in: Sparn, B. F. 2018. Laboratory Resources and Techniques to Evaluate Smart Home Technology (No. NREL/CP-5500-71696). National Renewable Energy Laboratory (NREL), Golden, CO (United States). https://www.nrel.gov/docs/fy18osti/71696.pdf and the experimental setup and validation of HVAC HIL platform can be found in: Ramaraj, S. and Sparn, B. 2022. Validation of HVAC Hardware-In-the-Loop Simulation for Advanced Control Strategies in Smart Homes (No. NREL/CP-5500-82562). National Renewable Energy Lab (NREL), Golden, CO (United States). https://www.nrel.gov/docs/fy22osti/82562.pdf. These experimental results can be used to validate how we currently model the cycling behavior of heat pumps and air conditioners. Additionally, many demand response programs implement heat pump and air conditioner control by changing the thermostat set point – these data may also be used to verify our models for heat pump and air conditioner demand response control are implemented correctly. The Test_Matrix file describes all the indoor and outdoor test conditions for heat pump and air conditioner and the file names of data sets include information about the test conditions. A wide range of outdoor air temperatures were chosen to accommodate summer and winter conditions. In addition to operating the HVAC equipment with different outdoor temperatures, we also operate the system with different indoor temperature set points to represent different grid signals or different operating conditions. For cooling conditions, the baseline set point is 72°F. To represent Load Up signals, the setpoint is changed to 68°F. The Load Shed set point is 76°F. For heating conditions, the baseline set point was assumed to be 68°F. The Load add set point is 72°F and the Load shed set point is 64°F. The starting indoor temperature for cooling conditions was set ~2°F above the indoor setpoint temperature so that the equipment turned on quickly. Similarly, the initial indoor temperature was set ~2°F lower than setpoint for heating mode tests to ensure that heating began quickly. The return air temperature was assumed to be equal to the indoor setpoint temperature in all cases. The experimental data are sampled at 1-second intervals. The data from ecobee thermostat at 5-minute interval are resampled and added to the corresponding file. The content of each data set is as follows: • T_Return (C): Measured return air temperature [C] • T_Return_SP (C): Return air temperature setpoint from E+ model, sent to HIL [C] • T_Supply (C): Measured supply air temperature at evaporator outlet [C] • T_Outdoor (C): Measured outdoor air temperature [C] • T_Outdoor_SP (C): Outdoor air temperature setpoint from weather file, sent to HIL [C] • T_Indoor (C): Measured indoor air temperature [C] • T_Indoor_SP (C): Indoor air temperature setpoint from E+ model, sent to HIL [C] • Outdoor Unit Power (W): Measured power of the outdoor unit [W] • Indoor Unit Power (W): Measured power of the indoor unit [W] • Evaporator Airflow Rate (CFM): Measured evaporator or indoor unit airflow rate sent to E+ model [CFM] • Cooling/Heating Capacity (kW): Calculated cooling/heating capacity sent to E+ model [kW] • T_SP_Thermostat (C): Thermostat cooling/heating setpoint temperature [C] • T_Indoor_Thermostat (C): Thermostat indoor air temperature [C]

24 POWER TRANSMISSION AND DISTRIBUTION