Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Advanced Materials & Manufacturing Technology (AMMT): Development of Additive Manufacturing Agnostic Process Parameter Procedure, 316H Stainless Steel Readiness Level Data Sets, and Machine Maintenance Plan

The University of California, Davis is involved in a project to deploy and enhance an artificial intelligence (AI) system for predicting and preventing plasma disruptions on the DIII D tokamak, under the funding from Department of Energy DE-SC0023500 (title: AI/Deep Learning FRNN Software for Prediction & Real-Time Control of DIII-D Plasma Control System (PCS)). The overarching goal is to demonstrate that real-time, AI-guided intervention can proactively modify the plasma state to avoid or mitigate disruptions—a critical challenge for the future of fusion energy.

36 MATERIALS SCIENCE

The ARM Precipitation Best Estimate (PrecipBE) Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s Precipitation Best-Estimate (PrecipBE) Value-Added Product integrates multiple precipitation datastreams, accounting for data quality and instrument limitations, to deliver comprehensive per-precipitation event properties alongside ancillary ARM data set data. PrecipBE bundles all valid surface rainfall samples into artificial intelligence (AI)-ready tabular and time-series formats, reporting bundle means and uncertainty ranges. This per-event structure provides an insightful and easy-to-use resource for researchers analyzing precipitation characteristics.

54 ENVIRONMENTAL SCIENCES

SPRUCE Vegetation Phenology in Experimental Plots from PhenoCam Imagery, 2015-2024

This data set consists of PhenoCam data from the SPRUCE experiment from the beginning of whole ecosystem warming (Hanson et al. 2017) in August 2015 through March 31 of 2025 (2015-08-24 to 2025-03-31), with start- and end-of-season phenological transition dates derived through the end of autumn 2024. Digital cameras, or phenocams, installed in each SPRUCE enclosure track seasonal variation in vegetation “greenness”, a proxy for vegetation phenology and associated physiological activity. Three separate regions of interest (ROIs) were defined for each camera field of view, corresponding to different vegetation types and demarcating (1) Picea trees (vegetation type EN, for evergreen needleleaf); (2) Larix trees (vegetation type DN, for deciduous needleleaf); and (3) the mixed shrub layer (vegetation type SH). This data set consists of three sets of data files: (1) 3-day summary product files: One file for each camera and each ROI (i.e. vegetation type), characterizing vegetation color at a 3-day time step. • Contains 36 files in *.csv format inside a compressed (*.zip) file. (2) Transition date file: Estimates “greenness rising” (spring) and “greenness falling” (autumn) transition dates derived from the smoothed daily green chromatic coordinate (GCC) values, for each camera and each ROI (i.e., vegetation type). • Contains one file in *.csv format. (3) Snow flag files: Indicate days with snow on trees or snow on ground for each experimental enclosure. • Contains two files in *.csv format, one for snow on trees and one for snow on ground. This data set consists of two sets of companion files: (1) Accompanying HTML files show the 90th quantiles of the mean GCC plotted together with transition dates for each vegetation type and plot. • Contains three files in HTML format, one for each vegetation type. • One additional file in HTML format with the transition dates plotted for each vegetation type, by year. (2) R files for processing PhenoCam files and flags. • Contains five files in R file(*.R) format and the components of the phenocamr package (Version 1.1.4) used for calculating transition dates for 2015-2024. These are contained in a compressed (*.zip) file. User Note: All imagery is posted in near-real time to the PhenoCam Project web page (https://phenocam.nau.edu), where it is publicly available. Scroll to “spruce” in the Gallery or link directly to the 29 SPRUCE cameras at https://tinyurl.com/sprucecams. This data set is based on the complete camera record from SPRUCE and supersedes all previously released PhenoCam datasets (see Related Data Sets). The estimated transition dates for previously released datasets may differ slightly (in most cases, by ±3 days or less), because following standard PhenoCam processing protocols (Richardson et al. 2018, Scientific Data), smoothing and interpolation, outlier removal, and transition date estimation are always conducted using the full data record.

54 ENVIRONMENTAL SCIENCES

Cambium 2024 Scenario Descriptions and Documentation

The National Renewable Energy Laboratory's (NREL's) Cambium data sets are annually released sets of simulated hourly data for a range of modeled futures of the U.S. electric sector with metrics designed to be useful for long-term decision- making. The 2024 Cambium data set is the fifth annual release. The data sets are a companion product to NREL's Standard Scenarios, which are likewise released annually and are a set of projections of how the U.S. electric sector could evolve across a suite of different potential futures, but covering more scenarios with less temporal granularity. Information about Cambium and related publications can be found at https://www.nrel.gov/analysis/cambium.html, and the Cambium data sets can be viewed and downloaded at https://scenarioviewer.nrel.gov/. In this documentation, we describe Cambium 2024's scenarios, define the metrics, and document the Cambium-specific methods for calculating those metrics.

24 POWER TRANSMISSION AND DISTRIBUTION

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Oak Ridge National Laboratory EAGLE-I TM : Modeling Electric Utility County Customers for Situational Awareness

During natural hazard events (hurricanes, wildfires, earthquakes, etc.) and recent man-made events (e.g., cyber attacks), the exchange of near real-time, spatially refined data within the response community is critical. The EAGLE-I$^{TM}$ platform is one tool that facilitates this data for decision makers within the energy sector. While much information can be collected and integrated into the system directly, other pertinent data must be augmented by other derived data products to enhance the information and allow for a consistent evaluation of on-the-ground conditions. One such data set that requires the addition of other derived data is the electric utility customer outage data that is aggregated to the county level within the EAGLE-I application. Without a county customer data set, outages can only be compared on total counts, which gives greater importance to higher population outages. Including an electric utility customer data set at the county level allows for these outage counts to be converted to percent outages and brings a consistent classification of outages and equal importance to all outages. To achieve this, several available data sets were combined and spatial disaggregation techniques were employed to model customer estimates at the county scale. This paper presents the approach to produce this data for the United States and lessons learned from working with these disparate data sets. Data validation is provided, where possible, and limitations of the model and possible improvements are discussed.

24 POWER TRANSMISSION AND DISTRIBUTION

Experimental data for damage mechanics simulation challenge

While there are many computational approaches for simulating damage in rock and other materials, few have been ground truth tested with either known experimental data or with blind data sets. Here, in this work, we present a bench-mark laboratory data set for a damage mechanics challenge to compare computational approaches on damage evolution in brittle-ductile materials. The samples were fabricated through additive manufacturing to produce repeatable specimens designed to fail in controlled ways. The failure was induced in the samples using a 3-point bending test to produce different Modes such as Mode I and mixed Modes including I-II, I-III and I-II-III Modes to generate a calibration data set and a blind challenge data set. Data collected included spatial and temporal measurements from traditional digital load–displacement sensors, 2D digital image correlation measurement to map surface deformations, 3D X-ray microscopy to ground-truth the crack-failure geometry, and laser profilometry to capture surface roughness. The data sets are available, on a data repository, to the community to advance computational models to improve our ability to predict damage in brittle-ductile materials.

3-point bending

Quality Ranking of Unary Chloride Salt Property Data Included in MSTDB-TP

Molten salt reactor developers rely on thermal property data to design, license and operate the reactors. The Molten Salt Thermal Database-Thermophysical Properties (MSTDB-TP) was established under the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program and is managed by Oak Ridge National Laboratory to serve as a single source of thermophysical property values measured for a wide variety of molten salt systems for use by researchers, molten salt reactor developers, and regulators. These properties include density, viscosity and thermal diffusivity and conductivity. Published measurements of molten salt properties are lacking for many salts of interest and the data that are available are often inconsistent. This creates a challenge for MSR developers when determining which property values to use when designing their reactors. It is the purpose of this work to apply a consistent ranking system to all data entries that indicates the quality of property values listed in the database. These rankings will be the technical basis for down-selections by the database developers and alert users about the quality of the available property values. MSTDB-TP collects all available property data and indicates preferred data sets or correlations. However, all available data sets are included in the database. Quality assessments and rankings are being applied to data in MSTDB-TP to provide an indication of the quality of each data set independent of consistency with other data. Previous reports detailed the ranking system that was followed and assessments of unary fluoride data sets. Documentation of the quality of data in MSTDB-TP was continued by reviewing and assessing all available sources of density, viscosity and thermal diffusivity or conductivity values for unary chloride salts in MSTDB-TP V3.0 using the same criteria.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

Towards an Improved Understanding of the Antarctic Coastal Zone and Its Contribution to Future Global Sea Level

Understanding the coastal zone of the Antarctic Ice Sheet (AIS), where it interacts with the Southern Ocean and warmer air masses, is crucial for predicting Antarctica's influence on the global climate and sea level. This region has multiple tipping mechanisms that could trigger large, rapid, and potentially irreversible changes in the AIS, the Southern Ocean and their global connections in the coming centuries. The AIS remains the largest source of uncertainty in future sea-level projections. Bed topography beneath the ice shelves and the coastal ice sheet is not yet well documented, and is a major source of this uncertainty. This review assesses current knowledge of the coastal zone and highlights methods to investigate it, including aerogeophysical surveys, ground- and ship-based measurements, satellite observations, and computer modeling. An ensemble analysis of published bed topography data sets identifies significant data gaps and their regional distribution, framed in the context of current ice-sheet behavior and potential instability. We propose scientific priorities and guidelines for future aerogeophysical surveys, advocating for a comprehensive, coordinated international effort to build a next-generation data set of Antarctic bed properties. Such an initiative would significantly advance understanding of the role of coastal processes in ice-sheet dynamics, reducing uncertainties in sea-level rise projections and improving predictions of future ocean and climate changes.

Kenichi Matsuoka

Satellite Imagery of PV Site Storm Damage

"This repository contains multiple data sets focused on visible damage to photovoltaic (PV) installations following extreme weather events such as hailstorms and hurricanes. Data sets are split into two categories: the first category, the ‘manually labeled’ data, was compiled by researchers manually, and contains manually identified PV sites exposed to storms. The second data set, the ‘aggregated’ data, is a compilation of the manually labeled PV sites and deep learning-identified PV sites. The hail damage data set focuses on post-storm PV damage following a September 24, 2023 hailstorm in Austin, TX, which caused over $600 million in damages in the Austin metro area. The hurricane damage data set focuses on post-storm PV damage following Hurricanes Irma and Maria in Puerto Rico and the US Virgin Islands. Hurricanes Irma and Maria were back-to-back category 5 hurricanes, which pummeled the Caribbean and southeastern United States in September 2017, causing an estimated $115.2 billion in damages."

14 SOLAR ENERGY

Long‐Term Large‐Scale Atmospheric Forcing Data From Three‐Dimensional Constrained Variational Analysis for the ARM SGP Site

Here, this study presents a long‐term three‐dimensional large‐scale forcing data set (VARANAL3D) derived from the three‐dimensional constrained variational analysis (3DCVA) method at the Atmospheric Radiation Measurement (ARM) program Southern Great Plains (SGP) site from 2004 to 2018. Building on the same input data sets as the conventional continuous forcing data set (VARANAL), VARANAL3D maintains overall consistency in domain‐averaged fields while introducing spatial variability, offering critical insights into the influence of mesoscale synoptic systems on cloud‐related processes. Evaluations are conducted across four cloud and precipitation regimes: Clear‐sky, Shallow‐clouds, Afternoon‐precipitation, and Nocturnal‐precipitation, presenting high consistency of the domain‐mean forcing data sets while emphasizing the role of subdomain forcing variability particularly in precipitating regimes. Single column model (SCM) simulations demonstrate that subdomain VARANAL3D forcing improves cloud and precipitation representation, with the ensemble outperforming domain‐mean forcing in three cloudy and precipitating regimes. Overall, these results highlight VARANAL3D's value for investigating the impacts of spatial variability of large‐scale forcing on atmospheric processes. The VARANAL3D data set provides new opportunities for evaluating model physics, advancing the development of scale‐aware parameterizations and deepening our understanding of cloud and precipitation dynamics.

Environmental sciences