Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

220EV Box Truck Telemetry Dataset

This dataset contains telemetry data for the 520EV refuse trucks accessed through a Veracity data platform as provided by Peterbilt. The dataset contains driving data on both electric trucks used by SWS. Data were recorded at an hourly resolution and contain energy use data while driving and idling, distance driven (miles), and driving speed (miles per hour). This dataset also contains charging data on the fast charger in kilowatt-hours for each hour.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

520EV Refuse Truck Telemetry Dataset

This dataset contains telemetry data for the 520EV refuse trucks accessed through a Veracity data platform as provided by Peterbilt. The dataset contains driving data on both electric trucks used by SWS. Data were recorded at an hourly resolution and contain energy use data while driving and idling, distance driven (miles), and driving speed (miles per hour). This dataset also contains charging data on the fast charger in kilowatt-hours for each hour.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset for Blueprinting Electrified Transit System Implementation

This dataset contains the figures and tabulated results generated from a system-level optimization study of transit fleet electrification planning. The dataset does not include executable modeling code required to reproduce the optimization. The dataset includes results for optimized charging infrastructure deployment by location and power level and service block assignments by fuel type, battery capacity selections, and distributed energy resource sizing. It also contains aggregated financial results, capital expenditures, operating cost summaries, net present cost comparisons across scenarios, and quantified air quality impacts. Results are structured to reflect multiple planning scenarios, including heuristic electrification plans, system-optimized configurations, and sensitivity cases with alternative objective weightings. The modeling was developed using publicly available General Transit Feed Specification data from Omnitrans and standardized modeling assumptions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

VECTOR Phase 1 Dataset: CAV Trajectory and Energy Consumption Records

This dataset contains benchmark experimental data from Phase 1 of the VECTOR project, focusing on the energy impact of CAV hardware components. The dataset includes vehicle trajectory data (speed and position) and corresponding energy consumption records collected from a CAV platform equipped with lidar, cameras, onboard computation units, and communication modules. The primary objective is to quantify the baseline energy consumption attributable to sensing and computing systems, independent of any advanced cooperative control strategies. During experiments, the leading vehicle followed a predetermined velocity profile, and the following CAV mirrored this trajectory using a basic car-following control to ensure consistent driving behavior. This setup enables a reliable benchmark for assessing the energy cost introduced by onboard CDA hardware (e.g., lidar and GPU-based processing). The dataset is essential for evaluating energy baselines and supports future comparative studies involving additional cooperative strategies. ![system img](system.png) ![vector img](vector.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset for Blueprinting Electrified Transit System Implementation

This dataset contains the figures and tabulated results generated from a system-level optimization study of transit fleet electrification planning. The dataset does not include executable modeling code required to reproduce the optimization. The dataset includes results for optimized charging infrastructure deployment by location and power level and service block assignments by fuel type, battery capacity selections, and distributed energy resource sizing. It also contains aggregated financial results, capital expenditures, operating cost summaries, net present cost comparisons across scenarios, and quantified air quality impacts. Results are structured to reflect multiple planning scenarios, including heuristic electrification plans, system-optimized configurations, and sensitivity cases with alternative objective weightings. The modeling was developed using publicly available General Transit Feed Specification data from Omnitrans and standardized modeling assumptions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Snow Depth Datasets for Snodgrass Catchment, Colorado, Water Year 2022-2023

This data package presents snow depths data from distributed temperature probes at 18 locations near Snodgrass catchment, Colorado. These data show that snow melt-out dates are approximately one or two weeks later under evergreen forests compared to other vegetation types even at the same elevation. These data were collected to understand how snowmelt heterogeneity impacts headwater hydrology, including streamflow and groundwater levels. They were also used to compare with process-based model simulations of snow depth to evaluate whether the model accurately represents snowmelt dynamics and their effects on headwater hydrology. Snow_DTPs_locations.csv includes all probes locations and their associated elevation and vegetation types. Snow_Depth_Snodgrass_WY2022_2023.csv includes processed snow depths datasets for Water Year (WY) 2022 and 2023. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. Several probes have recordings for WY 2021.

54 ENVIRONMENTAL SCIENCES↗

Dataset for Cruz-O'Byrne et al (2026): "Divergent biogeochemical responses in upland coastal forest soils to repeated flooding and shifts in water chemistry"

Hydrologic disturbances from accelerated sea-level rise and the increasing frequency and intensity of storms and tidal flooding are altering biogeochemical processes in upland coastal forests, transforming these ecosystems into wetlands. However, the initial effects of flooding on belowground biogeochemistry and the mechanisms driving greenhouse gas dynamics and soil organic matter stability during the early stages of this transition remain poorly understood. This dataset presents the results of a mesocosm experiment conducted in a controlled, highly instrumented laboratory environment, in which freshwater and brackish water pulses were applied to intact soil monoliths from a temperate upland coastal forest to examine how floodwater chemistry influences soil biogeochemistry and organo-mineral interactions. All data files are plain-text CSV (comma-separated value), and no special software is required to read them. Details about the content of each file are available in the document “Dataset_readme”. The dataset consists of the following data: • rcruzobyrne_moisture: Soil volumetric water content (VWC) • rcruzobyrne_GHG: Headspace greenhouse gas (GHG) concentration and fluxes • rcruzobyrne_methane_isotopes: Headspace methane isotope signature • rcruzobyrne_porewater: Porewater chemistry • rcruzobyrne_CDOM: Porewater colored dissolved organic matter (CDOM) • rcruzobyrne_FTIR: Soil Fourier-transform infrared (FTIR) spectroscopy Details of the experimental setup, data collection, and data analysis are provided in the manuscript by Cruz-O’Byrne et al (2026) Divergent biogeochemical responses in upland coastal forest soils to repeated flooding and shifts in water chemistry. Biogeochemistry. https://doi.org/10.1007/s10533-026-01340-0

EARTH SCIENCE > ATMOSPHERE > GREENHOUSE GAS↗

NEWTS Integrated Dataset (version 1.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v1.0 provides water researchers, community leaders, and regulators with a unified and standardized energy-related wastewater stream database. This resource is derived from 27 state and federal entities, and scientific publications, and contains more than 400,000 sample records, many of which also provide geospatial information. The dataset includes data for several different energy-related wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, and power plant wastewater. The NEWTS Integrated Dataset was built to support environmentally prudent decision-making, explore treatment opportunities, and identify potential critical mineral sources. A subset of this novel resource is also featured on NETL NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

abandoned mine drainage↗

Datasets and U-Net Model for "A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma"

This dataset has results and the model associated with the publication Ciulla et al., (2024). It contains a U-Net semantic segmentation model (unet_model.h5) and associated code implemented in tensorflow 2.0 for the model training and identification of oil and gas well symbols in USGS historical topographic maps (HTMC). Given a quadrangle map (7.5 minutes), downloadable at this url: https://ngmdb.usgs.gov/topoview/, and a list of coordinates of the documented wells present in the area, the model returns the coordinates of oil and gas symbols in the HTMC maps. For reproducibility of our workflow, we provide a sample map in California and the documented well locations for the entire State of California (CalGEM_AllWells_20231128.csv) downloaded from https://www.conservation.ca.gov/calgem/maps/Pages/GISMapping2.aspx. Additionally, the locations of 1,301 potential undocumented orphaned wells identified using our deep learning framework or the counties of Los Angeles and Kern in California, and Osage and Oklahoma in Oklahoma are provided in the file found_potential_UOWs.zip. The results of the visual inspection of satellite imagery in Osage County is in the file visible_potential_UOWs.zip. The dataset also includes a custom tool to validate the detected symbols in the HTMC maps (vetting_tool.py). More details about the methodology can be found in the associated paper: Ciulla, F., Santos, A., Jordan, P., Kneafsey, T., Biraud, S.C., and Varadharajan, C. (2024) A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma. Accepted for publication in Environmental Science and Technology. The geographical coordinates provided correspond to the locations of potential undocumented orphaned oil and gas wells (UOWs) extracted from historical maps. The actual presence of wells need to be confirmed with on-the-ground investigations. For your safety, do not attempt to visit or investigate these sites without appropriate safety training, proper equipment, and authorization from local authorities. Approaching these well sites without proper personal protective equipment (PPE) may pose significant health and safety risks. Oil and gas wells can emit hazardous gasses including methane, which is flammable, odorless and colorless, as well as hydrogen sulfide, which can be fatal even at low concentrations. Additionally, there may be unstable ground near the wellhead that may collapse around the wellbore. This dataset was prepared as an account of work sponsored by the United States Government. While this document is believed to contain correct information, neither the United States Government nor any agency thereof, nor the Regents of the University of California, nor any of their employees, makes any warranty, express or implied, or assumes any legal responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by its trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or the Regents of the University of California. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof or the Regents of the University of California.

Artificial Intelligence↗

NEWTS Integrated Dataset (version 2.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v2.0 provides water researchers, community leaders, regulators, and industry stakeholders with a unified and standardized energy-process wastewater chemistry database. This resource is derived from 39 state and federal entities, and scientific publications, and contains more than 700,000 sample records, many of which also provide geospatial information. The dataset includes chemistry data for several different energy-process wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, power plant wastewater, and geothermal fluids. The NEWTS Integrated Dataset was built to support prudent decision-making, characterization of potential critical mineral sources, and modeling of treatment and valorization options. A subset of this novel resource is also featured on the NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

AMD↗

Evaluating a Commercial Dynamic Line Rating Software with the National PMU Dataset

To accelerate the development of data-driven applications for power systems, the Department of Energy (DOE) supported the collection and curation of a synchrophasor dataset spanning two years of observations from transmission utilities across the US. This National PMU Dataset (NPDS) was anonymized and distributed to awardees of a DOE research grant under nondisclosure agreements (NDAs) but has also been retained at PNNL to enable further research. Agreements with data contributors prevent the data from being shared outside the organization. However, establishing a blind research validation methodology is envisioned to maximize the value proposition of the NPDS. In this validation strategy, researchers may share algorithms/software (potentially as executables to protect intellectual property) with PNNL, and PNNL will share feedback about the software’s performance on subsets of the NPDS. Such a blind methodology ensures that sensitive information about critical infrastructure remains protected, but the value of the NPDS can be extended to research beyond PNNL. Through iterative feedback, the algorithms may be tweaked to address real-world artifacts. As the NPDS data is temporally and geographically diverse, it may capture features absent in smaller datasets used during the development of the algorithm under test. This report presents lessons learned from applying the blind validation methodology to LineID™, a synchrophasor-based dynamic line rating software developed by Topolonet Corporation. Improvements made to the software through iterative feedback, limitations of the validation methodology, as well as how the limitations of the NPDS affected the evaluation process are discussed. Observations indicate that the proposed validation methodology can be valuable for evaluating other tools in the future.

97 MATHEMATICS AND COMPUTING↗

User Guide: A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations

This user guide supports a curated dataset of 71 meteor events recorded between 2006 and 2011 in Southwestern Ontario, Canada. Each event was simultaneously observed by ground-based optical cameras and an infrasound array, providing a rare opportunity to examine meteor trajectories and acoustic signals from the same atmospheric entry events. The dataset includes raw and processed optical data, meteor trajectories, photometric light curves, infrasound waveforms, and atmospheric specifications relevant for acoustic modeling. The archive is structured to support reproducible research in meteor physics, atmospheric acoustics, and shock wave analysis. It is organized following transparent file naming conventions and structured folders to facilitate scientific reuse, comparison, and integration across research domains. The dataset is freely available on Zenodo, doi: 10.5281/zenodo.15868512.

54 ENVIRONMENTAL SCIENCES↗

Development of an Unbiased Future Solar Dataset for Solar Resource Adequacy Research Over CONUS

A high-resolution, long-term solar dataset is essential for capturing the variability of solar energy resources and informing strategies to ensure grid reliability and resilience in systems with high levels of solar energy integration. This study focuses on generating unbiased, high-resolution projections of solar irradiance through a statistical downscaling framework, using Earth system model (ESM) simulations obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX). The National Solar Radiation Database (NSRDB) is used to calibrate statistical downscaling models. The newly developed dataset provides solar irradiance, surface air temperature, and surface wind speed at 4-km and hourly resolutions across the contiguous United States (CONUS), based on two future scenarios (RCP4.5 and RCP8.5). This study outlines key steps in developing the high-resolution future solar dataset, including (1) regridding ESM data to a common 20-km resolution grid, (2) correcting ESM biases using the NSRDB, and (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar projections. Preliminary results indicate that downscaled projections (4-km) captured reasonable spatial patterns when compared to observations across CONUS for four variables. On average across all pixels, 4-km daily-total GHI and DNI projections showed normalized bias (nBias) less than 1% and 6% for GHI and DNI against NSRDB, respectively (nBias less than 1% and 5% for daily-average surface air temperature and surface wind speed). In terms of long-term trend for GHI and DNI, there was no strong increasing or decreasing trend (when compared to surface air temperature), but it showed a very weak decreasing trend.

14 SOLAR ENERGY↗

Dataset, Code, and Models for Training Deep Learning Potentials for Low Temperature Plasma-Surface Interactions

This repository contains datasets, training scripts, and finished models, and test simulations used in the development of DeepREBO— a machine-learned interatomic potential trained to emulate the REBO2 empirical potential. The data was generated to study deep potential development for simulations of plasma-surface interactions. It uses an active learning framework, starting from a minimal dataset and iteratively expanding it. Included are those generated datasets, the trained models, and simulations used to evaluate the performance of the training process. This resource supports reproducibility and provides a reference framework for training deep potentials in plasma-surface interaction studies.

active learning↗

Synthetic Streamflow Datasets to Support Emulation of Water Allocations via LSTM

This archive is the data companion to the bonney_et-al_2026_erc metarepo which generates synthetic data, trains an LSTM model, and generates performance metrics on the trained model. While the generation of the synthetic data is fully reprodicible, it is a computationally expensive process. This data archive contains the synthetic datasets needed for training and testing an LSTM model and reproduction of figures and tables. In addition, supplemenatary data products generating and visualizing results is also included, such as geospatial data for the basin. Contents There are two high level directories: `WRAP_archive/` and `repo_data/`. The `WRAP_archive` directory contains compressed intermediate dataproducts from the dataset generation workflow (marked as "I_Dataset_Generation" in the metarepo). These data products are not required by any scripts in the metarepo, but they are archived as they are expensive to generate and may have useful information for other analyses. The `repo_data` directory contains the necessary data for reproducing the workflow in the metarepo and should be decompressed and moved into the top level of the metarepo. Additional details are provided in README.md.

drought↗

IM3 Projected U.S. Western Interconnection Grid Stress Dataset

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from GO WEST and TEP) for Integrated Multisector, Multiscale Modeling (IM3) Phase 2 simulations across eight different scenarios for the U.S. Western Interconnection through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States from a set of Thermodynamic Global Warming (TGW) simulations. These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 GO WEST is an open-source power grid modeling framework for the U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BAs) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of the U.S. Western Interconnection (ACTIVSg10k). Users can select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on the GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within the GO WEST framework. The TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. In order to use TEP model, users need to create scenarios with the GO WEST framework. Please refer to README file for a detailed description of the dataset including individual files and references.

Capacity Expansion Model↗

Legacy Survey of Space and Time Data Preview 2: deep_coadd dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the deep_coadd dataset type. These are the combination of multiple processed, calibrated, and background- subtracted images, for a patch of sky, for each of the six filters. This release contains 925,460 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 2: survey property dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the survey property dataset type. These are healSparse property maps for the survey. This release contains 84 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗