Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Information-entropy-driven generation of material-agnostic datasets for machine-learning interatomic potentials

In contrast to their empirical counterparts, machine-learning interatomic potentials (MLIAPs) promise to deliver near-quantum accuracy over broad regions of configuration space. However, due to their generic functional forms and extreme flexibility, they can catastrophically fail to capture the properties of novel, out-of-sample configurations, making the quality of the training set a determining factor, especially when investigating materials under extreme conditions. We propose a novel automated dataset generation method based on the maximization of the information entropy of the feature distribution, aiming at an extremely broad coverage of the configuration space in a way that is agnostic to the properties of specific target materials. The ability of the dataset to capture unique material properties is demonstrated on a range of unary materials, including elements with the FCC (Al), BCC (W), HCP (Be, Re and Os), graphite (C), and trigonal (Sb, Te) ground states. MLIAPs trained to this dataset are shown to be accurate over a range of application-relevant metrics, as well as extremely robust over very broad swaths of configurations space, even without dataset fine-tuning or hyper-parameter optimization, making the approach extremely attractive to rapidly and autonomously develop general-purpose MLIAPs suitable for simulations in extreme conditions.

36 MATERIALS SCIENCE

A unified ensemble soil moisture dataset across the continental United States

Abstract A unified ensemble soil moisture (SM) package has been developed over the Continental United States (CONUS). The data package includes 19 products from land surface models, remote sensing, reanalysis, and machine learning models. All datasets are unified to a 0.25-degree and monthly spatiotemporal resolution, providing a comprehensive view of surface SM dynamics. The statistical analysis of the datasets leverages the Koppen-Geiger Climate Classification to explore surface SM’s spatiotemporal variabilities. The extracted SM characteristics highlight distinct patterns, with the western CONUS showing larger coefficient of variation values and the eastern CONUS exhibiting higher SM values. Remote sensing datasets tend to be drier, while reanalysis products present wetter conditions. In-situ SM observations serve as the basis for wavelet power spectrum analyses to explain discrepancies in temporal scales across datasets facilitating daily SM records. This study provides a comprehensive soil moisture data package and an analysis framework that can be used for Earth system model evaluations and uncertainty quantification, quantifying drought impacts and land–atmosphere interactions and making recommendations for drought response planning.

54 ENVIRONMENTAL SCIENCES

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics

A global dataset of terrestrial biological nitrogen fixation

Biological nitrogen fixation (BNF) is the main natural source of new nitrogen inputs in terrestrial ecosystems, supporting terrestrial productivity, carbon uptake, and other Earth system processes. We assembled a comprehensive global dataset of field measurements of BNF in all major N-fixing niches across natural terrestrial biomes derived from the analysis of 376 BNF studies. The dataset comprises 32 variables, including site location, biome type, N-fixing niche, sampling year, quantification method, BNF rate (kg N ha −1 y −1 ), the percentage of nitrogen derived from the atmosphere (%N dfa ), N fixer or N-fixing substrate abundance, BNF rate per unit of N fixer abundance, and species identity. Overall, the dataset combines 1,207 BNF rates for trees, shrubs, herbs, soil, leaf litter, woody litter, dead wood, mosses, lichens, and biocrusts, 152 herb %N dfa values, 1,005 measurements of N fixer or N-fixing substrate abundance, and 762 BNF rates per unit of N fixer abundance for a total of 424 species across 66 countries. This dataset facilitates synthesis, meta-analysis, upscaling, and model benchmarking of BNF fluxes at multiple spatial scales.

Reis Ely, Carla R. [Oregon State Univ., Corvallis,

A dataset for understanding self-reported patterns influencing residential energy decisions

Household occupant behavior and decision-making dynamics substantially impact technology uptake and residential building energy performance. Although significant research underscores the importance of social science in energy studies, few public data with representative samples on household energy decision-making patterns are available. The dataset (UPGRADE-E: Understanding Patterns Guiding Residential Adoption and Decisions about Energy Efficiency) presents 9,919 responses from U.S. residents of single-family and small multifamily homes. Derived from a national-scale internet survey, the dataset contains 391 variables: demographics, building characteristics, home modifications, willingness to adopt new technologies, motivations for making changes, barriers, program participation, trusted information sources, and energy scenarios. Responses were validated via internal consistency checks and comparison with other U.S. national scale datasets. UPGRADE-E advances knowledge of household energy related decision-making, tying demographics, home modifications, and self-reported cognitive drivers together at a scale and breadth that has not been previously achieved. Policymakers and researchers at local, regional, and national levels may leverage this dataset to understand drivers influencing the adoption of key technologies in U.S. homes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

CAMELSH: A Large-Sample Hourly Hydrometeorological Dataset and Attributes at Watershed-Scale for CONUS

We present CAMELSH (Catchment Attributes and Hourly HydroMeteorology for Large-Sample Studies), the first large-sample hydrometeorological dataset at the hourly scale for the contiguous United States. CAMELSH intergrates hourly meteorological time series, catchment attributes and boundaries from GAGES-II and HydroATLAS for 9,008 catchments across diverse climatic, hydrological, and anthropogenic conditions. In addition, hourly streamflow time series is provided for 3,166 catchments. The dataset spans 45 years (1980–2024) with 11 meteorological variables from the NLDAS-2 forcing dataset, from which we compute nine climate indices related to precipitation, evapotranspiration, seasonality, and snow fraction. Additionally, CAMELSH includes two sets of catchment attributes: 439 from GAGES-II and 195 derived from HydroATLAS. These attributes include factors related to climate, geology, hydrology, river/stream morphology, landscape, nutrient, soil, topography, and anthropogenic influences. Developed in accordance with FAIR (Findability, Accessibility, Interoperability, and Reusability) principles, CAMELSH is the first large-sample dataset at an hourly timescale, supporting machine learning applications for short-term streamflow (flood) prediction and advancing data-driven hydrological research across multiple timescales.

54 ENVIRONMENTAL SCIENCES

Material Fracturing and Failure Simulation Datasets

Fracturing is a fundamental physics phenomena with broad relevance across multiple domains, ranging from infrastructure integrity, aerospace durability, reservoir production, and seismic events. We present a diverse dataset of simulated fracture evolution and material failure generated from two numerical solvers: the phase-field method and the combined finite-discrete element method (FDEM). These solvers differ in formulation, physical fidelity, and computational efficiency. The dataset includes five materials: PBX, anisotropic shale, tungsten, aluminum, and steel. For each, phase-field simulations span 400,000 cases: 200,000 under uniaxial tension and 200,000 under biaxial tension. The computationally expensive FDEM simulations include 90,000 split evenly among PBX, shale, and tungsten under uniaxial loading. All simulations begin with randomized initial fracture patterns. Each entry includes temporal data capturing fracture propagation dynamics. This comprehensive dataset is designed to support the development of foundational or surrogate machine learning approaches for predicting material failure. While no such models are introduced here, the dataset lays a robust foundation for advancing future research and innovation in these areas.

36 MATERIALS SCIENCE

Collection: TD-DFT and EOM-CCSD Calculations for the GDB-9-Ex Dataset

We present two datasets that contain quantum chemical electronic structure calculations for organic molecules from the GDB-9-Ex dataset. The “GDB-9-Ex_TD-DFT-PBE0” dataset contains calculations performed using the time-dependent density functional theory (TD-DFT) first principles method, and the “GDB-9-Ex_EOMCCSD” dataset contains calculations performed using the equation-of-motion coupled cluster (EOM-CCSD) method. Both types of calculations were performed using the ORCA software and provided ultraviolet-visible spectra with a high level of accuracy.

Mehta, Kshitij [Oak Ridge National Laboratory (ORN

Towards Diverse and Representative Global Pretraining Datasets for Remote Sensing Foundation Models

The design of a pretraining dataset is emerging as a critical component for the generality of foundation models. In the remote sensing realm, large volumes of imagery and benchmark datasets exist that can be leveraged to pretrain foundation models, however using this imagery in absence of a well-crafted sampling strategy is inefficient and has the potential to create biased and less generalizable models. Here, we provide a discussion and vision for the curation and assessment of pretraining datasets for remote sensing geospatial foundation models. We highlight the importance of geographic, temporal, and image acquisition diversity and review possible strategies to enable such diversity at global scale. In addition to these characteristics, support for various spatial-temporal pretext tasks within the dataset is also critical. Ultimately, our primary objective is to place emphasis on and draw attention to the data curation stage of the foundation model development pipeline. By doing so, we think it is possible to reduce biases of geospatial foundation models, as well as enable broader generalization to downstream remote sensing tasks and applications.

Arndt, Jacob

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan

RADAI: A Large-Scale Realistic Dataset for Radiation Detection Algorithm Development

Open, realistic datasets are essential for developing and benchmarking radiation detection algorithms, yet they remain scarce. The Radiological Anomaly Detection and Identification (RADAI) project was develop to create datasets that meet the training and testing needs for sophisticated radiation detection algorithms. The RADAI dataset is a large-scale synthetic resource that integrates high-fidelity Monte Carlo simulations with realistic urban scenarios to capture both background variability and source signatures. RADAI models construction-material NORM, people and vehicles, urban clutter, and dynamic environmental effects such as cosmic-ray and rain-induced transients, and they provide list-mode detector data with motion and response modeling suitable for algorithm training and evaluation. The RADAI project resulted in three publicly-released complementary datasets together with an online scoring portal for standardized performance assessment and an open software toolkit that supports data access, augmentation, model development, and evaluation. These resources enable reproducible comparisons across methods and promote rigorous studies at the scale required by contemporary machine learning. By grounding algorithm development in realistic, well-documented conditions, RADAI supports progress toward more robust detection, identification, and localization in complex urban environments.

Ghawaly, James M. [Division of Computer Science an

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)

BAMCensus (The Behavior and Advanced Mobility Census Dataset Aggregator) [SWR-25-120]

This software is a high-performance tool developed in Rust for downloading and processing large-scale geospatial datasets, specifically focusing on US Census data. It is designed to address scaling limitations found in existing tools, such as R's [tidycensus](https://walker-data.com/tidycensus/), by providing performant streaming dataset JOIN operations between various US Census datasets (like ACS and LEHD) and their corresponding geometries stored on the TIGER/Lines web server. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. Its primary motivation stems from the need for a high-performance solution to combine spatial datasets with graph traversals within the context of mobility analysis tooling being developed at NREL's Behavior and Advanced Mobility (BAM) group.

Fitzgerald, Robert [National Renewable Energy Labo

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION

2024 Buildings Technology Baseline: Dataset Documentation

The Buildings Technology Baseline is a curated and regularly updated dataset of current and projected performance, retail, and installed price data for all major building energy technologies needed to enable cost/benefit analyses. Building technology analyses require an up-to-date understanding of installation costs and cost-effectiveness of key building energy efficiency technologies. The dataset was assembled by Guidehouse during fiscal year 2024. Data was gathered from the 2024 National Residential Efficiency Measures Database (NREMDB), the 2023 Energy Information Administration Updated Buildings Sector Appliance and Equipment Costs and Efficiencies ("EIA Building Data Report"), DOE Lighting Market Model, the 2023 RSMeans database, and the 2020 Grid-Interactive Efficient Building Technology Cost, Performance, and Lifetime Characteristics ("GEB Data Report"), Lawrence Berkeley National Laboratory, various literature, as well as new data from online retailers, stakeholder interviews, and contractor databases in 2023 and 2024. The dataset has been reviewed by subject matter experts at NREL and DOE. The 2024 dataset release is intended to be a starting point for interested users to provide feedback. This database is not intended to provide specific cost estimates for a specific project. The cost estimates do not include any rebates or tax incentives that may be available for the measures. Rather, it is meant to help determine which measures may be more cost-effective. The National Renewable Energy Laboratory (NREL) makes every effort to ensure accuracy of the data; however, NREL does not assume any legal liability or responsibility for the accuracy or completeness of the information.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Development of A High-Resolution Dataset for Solar Resource Adequacy Studies

High-resolution, long-term solar dataset is essential for characterizing the variability of solar energy resources and for informing strategies that ensure grid reliability and resilience in grid systems with high levels of solar energy integration. We present the development of a new 4-km, hourly Earth system dataset for the contiguous United States (CONUS), using a statistical downscaling approach that integrates the National Solar Radiation Database (NSRDB) with regional Earth system model projections. The new high-resolution Earth system dataset includes key variables - GHI, DNI, DHI, surface air temperature, and wind speed - under two future scenarios. Preliminary results show a reasonable agreement with NSRDB observations, with nBias less than 1% for GHI across CONUS. The dataset is expected to support in-depth analyses of extreme weather impacts and provide input to resource adequacy for future energy systems with diverse generation sources.

14 SOLAR ENERGY

Algorithmically detected rain-on-snow flood events in different climate datasets: a case study of the Susquehanna River basin

Abstract. Rain-on-snow (RoS) events in regions of ephemeral snowpack – such as the northeastern United States – can be key drivers of cool-season flooding. We describe an automated algorithm for detecting basin-scale RoS events in gridded climate data by generating an area-averaged time series and then searching for periods of concurrent precipitation, surface runoff, and snowmelt exceeding predefined thresholds. When evaluated using historical data over the Susquehanna River basin (SRB), the technique credibly finds RoS events in published literature and flags events that are followed by anomalously high streamflow as measured by gauge data along the river. When comparing four different datasets representing the same 21-year period, we find large differences in RoS event magnitude and frequency, primarily driven by differences in estimated surface runoff and snowmelt. Using dataset-specific thresholds improves agreement between datasets but does not account for all discrepancies. We show that factors such as meteorological forcing and coupling frequency, as well as choice of land surface model, play roles in how data products capture these compound extremes and suggest care is to be taken when climate datasets are used by stakeholders for operational decision-making.

54 ENVIRONMENTAL SCIENCES

Farm Practice Typologies as a Strategy for Management-Relevant Land Use and Land Cover Mapping in the Great Lakes Region (Version 1) [Dataset]

Dataset overview and development This dataset provides spatially explicit agricultural land-use and land-management typologies developed for the Great Lakes Region (GLR) at the farm-parcel level. The typologies were designed to characterize not only the land-use and land-cover (LULC) associated with individual agricultural farm parcels, but also the land-management practices (LMPs), including irrigation, tile drainage, and conservation easements, occurring within those parcels and how these characteristics change through time. The dataset contains four related typology products: Annual integrated typology – describes the combined LULC and land-management characteristics for each farm parcel for individual years. LULC transition typology – describes the temporal pattern of LULC change for each farm parcel across the study period (2008-2023). LMP trend typology – describes the temporal pattern in the occurrence of LMPs for each farm parcel across the study period. Multi-year integrated typology – combines the LULC transition typology and LMP trend typology to provide an integrated characterization of long-term land-use and management patterns. Purpose of the dataset The purpose of these products is to provide a management-relevant integrated and consistent framework for evaluating the spatial and temporal organization of agricultural landscapes across the GLR. The resulting typologies can: support landscape-scale environmental and land-use analysis; provide spatial information relevant to land-management strategies, conservation planning, policy development, and program evaluation; characterize spatial patterns of agricultural land use and management; examine changes in agricultural landscapes through time; and identify persistent, transitional, and changing agricultural systems. Please refer to the README file provided in Files for more details.

Agriculture