Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David↗

The high explosives & affected targets (HEAT) dataset

Artificial Intelligence (AI) surrogate models offer a computationally efficient alternative to full-physics simulations, yet no existing datasets are publicly available for training, testing, and validation of machine learning models of the dynamics of high-explosive driven shocks through multiple materials. Shock propagation through materials is a computationally challenging problem because simulations must include material-specific equations of state (EOS) along with descriptions of other physical processes such as plastic deformation, phase change, damage processes, fluid instabilities, and multi-material interactions. Shocks are typically initiated by high-velocity impacts or explosive loading. The latter case necessitates the addition of models of reactive materials to represent high-explosive (HE) detonation. Here, to address the lack of an expansive dataset for multi-material shock propagation in the AI/ML community, we present the High-Explosives and Affected Targets (HEAT) Dataset. HEAT is a physics-rich collection of two-dimensional, cylindrically symmetric, simulations generated using an Eulerian, multi-material, shock-propagation code developed at Los Alamos National Laboratory. The dataset includes two partitions: (1) the expanding shock-cylinder (CYL) simulations, Figs. 1, and (2) the Perturbed Layered Interface (PLI) simulations, Fig. 2. Entries in both partitions consist of time series of arrays of thermodynamic fields (pressure, density, and temperature), kinematic fields (position and velocity), and additional fields that depend on thermodynamic and/or kinematic fields (e.g., material stress). Materials in the CYL partition include solids (aluminium, copper, depleted uranium, stainless steel, tantalum, and a generic polymer), a liquid (water), gases (air, nitrogen), and a generic detonating material (high explosive, HE). The PLI partition spans a highly varying geometry but consists of fixed materials across entries: Copper, aluminium, stainless steel, generic polymer, and generic HE. HEAT captures critical phenomena such as momentum transfer, shock propagation, plastic deformation, and thermal effects, making HEAT a valuable benchmark for development of AI/ML emulation of multi-material shock propagation.

36 MATERIALS SCIENCE↗

Systematic Evaluation of Atmospheric Forcing, Surface Datasets, and Mesh Effects on Kilometer-Scale Land Surface and River Modeling

Earth system models are advancing toward kilometer-scale resolution to capture local climate impacts and extremes. High-resolution land and river modeling depends on multiple factors, including mesh, surface datasets, and atmospheric forcing, but their relative effects at kilometer scales remain unquantified. We evaluated five Energy Exascale Earth System Model land and river configurations over the Mid-Atlantic region using two mesh (1/8° structured versus variable-resolution unstructured mesh), two surface datasets (default versus newly developed), and three atmospheric forcings (NLDAS2, MSWX, GSWP). Evaluation against satellite, reanalysis, and in situ benchmarks across water, energy, and carbon cycles quantifies how these factors affect model performance. Forcing selection produces the largest bias reductions (12-99% across variables), followed by surface datasets (7-75%) and mesh (up to 21%). Forcing effects vary by variable, with MSWX reducing biases for snow water equivalent, evapotranspiration, albedo, temperature, and gross primary productivity, GSWP for snow cover and runoff, and NLDAS for soil moisture and streamflow. The use of newly developed surface datasets improves gross primary productivity (58% bias reduction) and evapotranspiration but increase soil moisture and albedo biases due to current modeling limitations. Variable-resolution unstructured mesh improves the simulation of small-basin streamflow through better capturing drainage networks, though mesh minimally affects other land variables. These findings provide important guidance for high-resolution modeling development and actionable science.

Land and River modeling↗

Shared micromobility as a first- and last-mile transit solution? Spatiotemporal insights from a novel dataset

The first- and last-mile (FM/LM) problem is a major deterrent to public transit use. With the rise of shared micromobility options such as shared e-scooters in recent years, there is a growing interest in understanding their potential to serve as a last-mile transit solution. However, empirical data regarding the integrated use of shared micromobility and public transit have been limited so far. As a result, much is unknown regarding the spatiotemporal patterns and characteristics of shared micromobility trips serving as an FM/LM connection to transit. Here, this paper addresses these knowledge gaps by leveraging a novel dataset (i.e., the Spin post-ride survey dataset) that records thousands of transit-connecting shared e-scooter trips in Washington DC. Specifically, we used the dataset to reveal the spatiotemporal patterns of transit-connecting shared e-scooter trips in Washington DC, resulting in some major policy insights regarding the integral use of shared e-scooters and public transit. We further leveraged the dataset to validate if and to what extent a commonly applied buffer-zone approach can infer FM/LM micromobility trips accurately. Statistical tests showed that the actual FM/LM Spin e-scooter trips differ from inferred FM/LM Spin e-scooter trips in both spatial and temporal dimensions. This indicates that the common practice of inferring FM/LM micromobility trips with a buffer-zone approach can lead to inaccurate estimates of transit-connecting micromobility trips.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A global urban heat island intensity dataset: Generation, comparison, and analysis

The urban heat island (UHI) effect, a phenomenon of local warming over urban areas, is the most well-known impact of urbanization on climate. Globally consistent estimates of the UHI intensity (UHII) are crucial for examining this phenomenon across time and space. However, publicly available UHII datasets are limited and have several constraints: (1) they are for clear-sky surface UHII, not all-sky surface UHII and canopy (air temperature) UHII; (2) the estimation methods often neglect anthropogenic disturbance, introducing uncertainties in the estimated UHII. To address these issues, this study proposes a new dynamic equal-area (DEA) method that can minimize the influence of various confounding factors on UHII estimates through a dynamic cyclic process. Utilizing the DEA method and leveraging various gridded temperature data, we develop a global-scale (>10,000 cities), long-term (over 20 years by month), and multi-faceted (clear-sky surface, all-sky surface, and canopy) UHII dataset. Further, based on these estimates, we provide a comprehensive analysis of the UHII and its trends in global cities. The UHII is found to be greater than zero in >80% of cities, with global annual average magnitudes around 1.0 °C (day) and 0.8 °C (night) for surface UHII, and close to 0.5 °C for canopy UHII. Furthermore, an interannual upward trend in UHII is observed in >60% of cities, with global annual average trends exceeding 0.1 °C/decade (day) and over 0.06 °C/decade (night) for surface UHII, and slightly surpassing 0.03 °C/decade for canopy UHII. Notably, there exists a positive correlation between the magnitude and trend of UHII, suggesting that cities with stronger UHII tend to experience faster growth in UHII. Additionally, discrepancies in UHII are found between different temperature data, stemming not only from distinctions in data types (surface or air temperature) but also from differences in data acquisition times (Terra or Aqua), weather conditions (clear-sky or all-sky), and processing methodologies (with or without gap filling). Overall, our proposed method, dataset, and analysis results have the potential to provide valuable insights for future urban climate studies. The UHII dataset is publicly available at https://doi.org/10.6084/m9.figshare.24821538.

54 ENVIRONMENTAL SCIENCES↗

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics↗

Creation of Polymer Datasets with Targeted Backbones for Screening of High-Performance Membranes for Gas Separation

A simple approach was developed to computationally construct a polymer dataset by combining simplified molecular-input line-entry system (SMILES) strings of a targeted polymer backbone and a variety of molecular fragments. This method was used to create 14 polymer datasets by combining seven polymer backbones and molecules from two large molecular datasets (MOSES and QM9). Polymer backbones that were studied include four polydimethylsiloxane (PDMS) based backbones, poly(ethylene oxide) (PEO), poly(allyl glycidyl ether) (PAGE), and polyphosphazene (PPZ). The generated polymer datasets can be used for various cheminformatics tasks, including high-throughput screening for gas permeability and selectivity. This study utilized machine learning (ML) models to screen the polymers for CO2/CH4 and CO2/N2 gas separation using membranes. Several polymers of interest were identified. Here the results highlight that employing an ML model fitted to polymer selectivities leads to higher accuracy in predicting polymer selectivity compared to using the ratio of predicted permeabilities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Information-entropy-driven generation of material-agnostic datasets for machine-learning interatomic potentials

In contrast to their empirical counterparts, machine-learning interatomic potentials (MLIAPs) promise to deliver near-quantum accuracy over broad regions of configuration space. However, due to their generic functional forms and extreme flexibility, they can catastrophically fail to capture the properties of novel, out-of-sample configurations, making the quality of the training set a determining factor, especially when investigating materials under extreme conditions. We propose a novel automated dataset generation method based on the maximization of the information entropy of the feature distribution, aiming at an extremely broad coverage of the configuration space in a way that is agnostic to the properties of specific target materials. The ability of the dataset to capture unique material properties is demonstrated on a range of unary materials, including elements with the FCC (Al), BCC (W), HCP (Be, Re and Os), graphite (C), and trigonal (Sb, Te) ground states. MLIAPs trained to this dataset are shown to be accurate over a range of application-relevant metrics, as well as extremely robust over very broad swaths of configurations space, even without dataset fine-tuning or hyper-parameter optimization, making the approach extremely attractive to rapidly and autonomously develop general-purpose MLIAPs suitable for simulations in extreme conditions.

36 MATERIALS SCIENCE↗

A three-year dataset supporting research on building energy management and occupancy analytics

Abstract This paper presents the curation of a monitored dataset from an office building constructed in 2015 in Berkeley, California. The dataset includes whole-building and end-use energy consumption, HVAC system operating conditions, indoor and outdoor environmental parameters, as well as occupant counts. The data were collected during a period of three years from more than 300 sensors and meters on two office floors (each 2,325 m 2 ) of the building. A three-step data curation strategy is applied to transform the raw data into research-grade data: (1) cleaning the raw data to detect and adjust the outlier values and fill the data gaps; (2) creating the metadata model of the building systems and data points using the Brick schema; and (3) representing the metadata of the dataset using a semantic JSON schema. This dataset can be used in various applications—building energy benchmarking, load shape analysis, energy prediction, occupancy prediction and analytics, and HVAC controls—to improve the understanding and efficiency of building operations for reducing energy use, energy costs, and carbon emissions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Molecular structural dataset of lignin macromolecule elucidating experimental structural compositions

Abstract Lignin is one of the most abundant biopolymers in nature and has great potential to be transformed into high-value chemicals. However, the limited availability of molecular structure data hinders its potential industrial applications. Herein, we present the Lignin Structural (LGS) Dataset that includes the molecular structure of milled wood lignin focusing on two major monomeric units (coniferyl and syringyl), and the six most common interunit linkages (phenylpropane β-aryl ether, resinol, phenylcoumaran, biphenyl, dibenzodioxocin, and diaryl ether). The dataset constitutes a unique resource that covers a part of lignin’s chemical space characterized by polymer chains with lengths in the range of 3 to 25 monomer units. Structural data were generated using a sequence-controlled polymer generation approach that was calibrated to match experimental lignin properties. The LGS dataset includes 60 K newly generated lignin structures that match with high accuracy (~90%) the experimentally determined structural compositions available in the literature. The LGS dataset is a valuable resource to advance lignin chemistry research, including computational simulation approaches and predictive modelling.

Scientific data↗

A large expert-curated cryo-EM image dataset for machine learning protein particle picking

Cryo-electron microscopy (cryo-EM) is a powerful technique for determining the structures of biological macromolecular complexes. Picking single-protein particles from cryo-EM micrographs is a crucial step in reconstructing protein structures. However, the widely used template-based particle picking process is labor-intensive and time-consuming. Though machine learning and artificial intelligence (AI) based particle picking can potentially automate the process, its development is hindered by lack of large, high-quality labelled training data. To address this bottleneck, we present CryoPPP, a large, diverse, expert-curated cryo-EM image dataset for protein particle picking and analysis. It consists of labelled cryo-EM micrographs (images) of 34 representative protein datasets selected from the Electron Microscopy Public Image Archive (EMPIAR). The dataset is 2.6 terabytes and includes 9,893 high-resolution micrographs with labelled protein particle coordinates. The labelling process was rigorously validated through 2D particle class validation and 3D density map validation with the gold standard. The dataset is expected to greatly facilitate the development of both AI and classical methods for automated cryo-EM protein particle picking.

42 ENGINEERING↗

The Arctic Plant Aboveground Biomass Synthesis Dataset

Plant biomass is a fundamental ecosystem attribute that is sensitive to rapid climatic changes occurring in the Arctic. Nevertheless, measuring plant biomass in the Arctic is logistically challenging and resource intensive. Lack of accessible field data hinders efforts to understand the amount, composition, distribution, and changes in plant biomass in these northern ecosystems. Here, we present The Arctic plant aboveground biomass synthesis dataset, which includes field measurements of lichen, bryophyte, herb, shrub, and/or tree aboveground biomass (g m -2 ) on 2,327 sample plots from 636 field sites in seven countries. We created the synthesis dataset by assembling and harmonizing 32 individual datasets. Aboveground biomass was primarily quantified by harvesting sample plots during mid- to late-summer, though tree and often tall shrub biomass were quantified using surveys and allometric models. Each biomass measurement is associated with metadata including sample date, location, method, data source, and other information. This unique dataset can be leveraged to monitor, map, and model plant biomass across the rapidly warming Arctic.

54 ENVIRONMENTAL SCIENCES↗

A unified ensemble soil moisture dataset across the continental United States

Abstract A unified ensemble soil moisture (SM) package has been developed over the Continental United States (CONUS). The data package includes 19 products from land surface models, remote sensing, reanalysis, and machine learning models. All datasets are unified to a 0.25-degree and monthly spatiotemporal resolution, providing a comprehensive view of surface SM dynamics. The statistical analysis of the datasets leverages the Koppen-Geiger Climate Classification to explore surface SM’s spatiotemporal variabilities. The extracted SM characteristics highlight distinct patterns, with the western CONUS showing larger coefficient of variation values and the eastern CONUS exhibiting higher SM values. Remote sensing datasets tend to be drier, while reanalysis products present wetter conditions. In-situ SM observations serve as the basis for wavelet power spectrum analyses to explain discrepancies in temporal scales across datasets facilitating daily SM records. This study provides a comprehensive soil moisture data package and an analysis framework that can be used for Earth system model evaluations and uncertainty quantification, quantifying drought impacts and land–atmosphere interactions and making recommendations for drought response planning.

54 ENVIRONMENTAL SCIENCES↗

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics↗

A global dataset of terrestrial biological nitrogen fixation

Biological nitrogen fixation (BNF) is the main natural source of new nitrogen inputs in terrestrial ecosystems, supporting terrestrial productivity, carbon uptake, and other Earth system processes. We assembled a comprehensive global dataset of field measurements of BNF in all major N-fixing niches across natural terrestrial biomes derived from the analysis of 376 BNF studies. The dataset comprises 32 variables, including site location, biome type, N-fixing niche, sampling year, quantification method, BNF rate (kg N ha −1 y −1 ), the percentage of nitrogen derived from the atmosphere (%N dfa ), N fixer or N-fixing substrate abundance, BNF rate per unit of N fixer abundance, and species identity. Overall, the dataset combines 1,207 BNF rates for trees, shrubs, herbs, soil, leaf litter, woody litter, dead wood, mosses, lichens, and biocrusts, 152 herb %N dfa values, 1,005 measurements of N fixer or N-fixing substrate abundance, and 762 BNF rates per unit of N fixer abundance for a total of 424 species across 66 countries. This dataset facilitates synthesis, meta-analysis, upscaling, and model benchmarking of BNF fluxes at multiple spatial scales.

Reis Ely, Carla R. [Oregon State Univ., Corvallis,↗

A dataset for understanding self-reported patterns influencing residential energy decisions

Household occupant behavior and decision-making dynamics substantially impact technology uptake and residential building energy performance. Although significant research underscores the importance of social science in energy studies, few public data with representative samples on household energy decision-making patterns are available. The dataset (UPGRADE-E: Understanding Patterns Guiding Residential Adoption and Decisions about Energy Efficiency) presents 9,919 responses from U.S. residents of single-family and small multifamily homes. Derived from a national-scale internet survey, the dataset contains 391 variables: demographics, building characteristics, home modifications, willingness to adopt new technologies, motivations for making changes, barriers, program participation, trusted information sources, and energy scenarios. Responses were validated via internal consistency checks and comparison with other U.S. national scale datasets. UPGRADE-E advances knowledge of household energy related decision-making, tying demographics, home modifications, and self-reported cognitive drivers together at a scale and breadth that has not been previously achieved. Policymakers and researchers at local, regional, and national levels may leverage this dataset to understand drivers influencing the adoption of key technologies in U.S. homes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

CAMELSH: A Large-Sample Hourly Hydrometeorological Dataset and Attributes at Watershed-Scale for CONUS

We present CAMELSH (Catchment Attributes and Hourly HydroMeteorology for Large-Sample Studies), the first large-sample hydrometeorological dataset at the hourly scale for the contiguous United States. CAMELSH intergrates hourly meteorological time series, catchment attributes and boundaries from GAGES-II and HydroATLAS for 9,008 catchments across diverse climatic, hydrological, and anthropogenic conditions. In addition, hourly streamflow time series is provided for 3,166 catchments. The dataset spans 45 years (1980–2024) with 11 meteorological variables from the NLDAS-2 forcing dataset, from which we compute nine climate indices related to precipitation, evapotranspiration, seasonality, and snow fraction. Additionally, CAMELSH includes two sets of catchment attributes: 439 from GAGES-II and 195 derived from HydroATLAS. These attributes include factors related to climate, geology, hydrology, river/stream morphology, landscape, nutrient, soil, topography, and anthropogenic influences. Developed in accordance with FAIR (Findability, Accessibility, Interoperability, and Reusability) principles, CAMELSH is the first large-sample dataset at an hourly timescale, supporting machine learning applications for short-term streamflow (flood) prediction and advancing data-driven hydrological research across multiple timescales.

54 ENVIRONMENTAL SCIENCES↗

Material Fracturing and Failure Simulation Datasets

Fracturing is a fundamental physics phenomena with broad relevance across multiple domains, ranging from infrastructure integrity, aerospace durability, reservoir production, and seismic events. We present a diverse dataset of simulated fracture evolution and material failure generated from two numerical solvers: the phase-field method and the combined finite-discrete element method (FDEM). These solvers differ in formulation, physical fidelity, and computational efficiency. The dataset includes five materials: PBX, anisotropic shale, tungsten, aluminum, and steel. For each, phase-field simulations span 400,000 cases: 200,000 under uniaxial tension and 200,000 under biaxial tension. The computationally expensive FDEM simulations include 90,000 split evenly among PBX, shale, and tungsten under uniaxial loading. All simulations begin with randomized initial fracture patterns. Each entry includes temporal data capturing fracture propagation dynamics. This comprehensive dataset is designed to support the development of foundational or surrogate machine learning approaches for predicting material failure. While no such models are introduced here, the dataset lays a robust foundation for advancing future research and innovation in these areas.

36 MATERIALS SCIENCE↗