Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

NETL RDE Image Classification Dataset 2025 - 14 Classes

Dataset including high-speed down-axis RDE images used for updated image classification study. This dataset includes 180,000 images with 14 classifications: 1CW, 1CCW, 2CW, 2CCW, 3CW, 3CCW, Deflagration, 4CW, 4CCW, 5CW, and 5CCW. Images are cropped to center annulus, and resized to 301x301 pixels. Images are filtered using the AFRL Beta correction factor.

Dataset↗

NETL RDE Image Classification Dataset 2020 - 10 Classes

Dataset including high-speed down-axis RDE images used for updated image classification study. This dataset includes 100,000 images with 10 classifications: 1CW, 1CCW, 2CW, 2CCW, 3CW, 3CCW, and Deflagration. Images are cropped to center annulus, and resized to 301x301 pixels. Images are filtered using the AFRL Beta correction factor.

AS↗

FlowDash Geothermal Energy Enhancer: Where is Next Geothermal Resource? Machine Learning + Multiple Datasets => Geothermal Exploration Indication?

This is the presentation delivered at the 2025 GEODE Datathon competition. GEODE is a consortium of experts that addresses technology and knowledge gaps in geothermal energy, leveraging technology and best practices from the oil and gas industry. NETL team was awarded the 1st place in the engineering track. 2025 GEODE Datathon had a total of 42 teams from top universities and several major industrial companies. This awarded work is founded on a robust idea and innovative approach that uses machine learning coupled to multiple datasets to visualize geothermal “sweet” spots/indications in Great Basin based on the data provided from the GEODE Datathon. The use case also leveraged other datasets and demonstrated insightful and valuable indications for geothermal exploration.

Geothermal energy, Machine learning, Multiple Data↗

A Forward-Looking Dataset of EV Managed Charging Resource and Costs

This presentation summarizes a high-resolution, forward-looking dataset of EV adoption, EV charging, and managed charging resource. Vehicle-level data are grounded in current adoption and charging patterns, and ~200,000 real-world vehicle-weeks of travel data covering all on-road segments (i.e., light-duty, transit and school buses, local, regional and long-haul medium- and heavy-duty). The data, which include multiple charging profiles per vehicle to bound flexibility, are then processed and aggregated to describe baseline charging and charge management resource by county, hour, year, scenario, and vehicle type. Coupled with one of four scenarios of how EV managed charging costs might evolve over time, the dataset enables a power sector capacity expansion model to select cost-optimal quantities of EV managed charging and supply-side resources to reliably satisfy demand. Five integration strategies: Baseline, Daytime and Flat (passive), Flex (active), and Stress (anti-strategy), illustrate how baseline charging and flexibility potential changes with EVSE build-out and charging preferences.

33 ADVANCED PROPULSION SYSTEMS↗

The Foundational Industry Energy Dataset: Unit-level Characterization and Derived Energy Estimates for Industrial Facilities in 2017

The Foundational Industry Energy Dataset (FIED) addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization by facility. Each facility is identified by a unique registryID, based on the U.S. Environmental Protection Agency (EPA) Facility Registry Service, and includes its coordinates and other geographic identifiers. Energy-using units are characterized by design capacity, as well as their estimated energy use, greenhouse gas emissions, and physical throughput using 2017 data from the EPA's National Emissions Inventory and Greenhouse Gas Reporting Program. An overview of the derivation methods is provided in a separate technical report which will be linked after publication. The Python code used to compile the dataset is available in a GitHub repository. An updated 2020 version is under development.

Array↗

Historical Bolide Infrasound Dataset (1960–1972)

We present the first fully curated, publicly accessible archive of infrasonic records from ten large bolide events documented by the U.S. Air Force Technical Applications Center’s global microbarometer network between 1960 and 1972. Captured on analog strip-chart paper, these waveforms predate modern digital arrays and space-based sensors, making them a unique window on meteoroid activity in the mid-twentieth century. Prior studies drew important scientific conclusions from the records but released only limited artifacts, chiefly period–amplitude tables and unprocessed scans, leaving the underlying data inaccessible for independent study. The present release transforms those limited excerpts into a research-ready resource. By capturing ten large events in the mid-20th century, the dataset constitutes a critical reference point for assessing bolide activity before the advent of modern space-based and digital ground-based monitoring. The multi-year coverage and worldwide distribution of events provide a valuable reference for comparing past and more recent detections, facilitating assessments of long-term flux and the dynamics of acoustic wave propagation in Earth’s atmosphere. The dataset’s availability in a consolidated format ensures straightforward access to waveforms and derived measurements, supporting a wide range of scientific inquiries into bolide physics and infrasound monitoring. By preserving these historical acoustic observations, the collection maintains a significant record of mid-20th-century meteoroid entries. It thereby establishes a basis for further refinement of impact hazard evaluations, contributes to historical continuity in atmospheric observation, and enriches the study of meteoroid-generated infrasound signals on a global scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗

Dataset of simulated vibrational density of states and X-ray diffraction profiles of mechanically deformed and disordered atomic structures in Gold, Iron, Magnesium, and Silicon

This dataset is comprised of a library of atomistic structure files and corresponding X-ray diffraction (XRD) profiles and vibrational density of states (VDoS) profiles for bulk single crystal silicon (Si), gold (Au), magnesium (Mg), and iron (Fe) with and without disorder introduced into the atomic structure and with and without mechanical loading. Included with the atomistic structure files are descriptor files that measure the stress state, phase fractions, and dislocation content of the microstructures. All data was generated via molecular dynamics or molecular statics simulations using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) code. This dataset can inform the understanding of how local or global changes to a materials microstructure can alter their spectroscopic and diffraction behavior across a variety of initial structure types (cubic diamond, face-centered cubic (FCC), hexagonal close-packed (HCP), and body-centered cubic (BCC) for Si, Au, Mg, and Fe, respectively) and overlapping changes to the microstructure (i.e., both disorder insertion and mechanical loading).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Re-evaluating probable maximum precipitation estimates: sensitivity to transposition domains and storm rotation using modern datasets

This study examines the sensitivity of Probable Maximum Precipitation (PMP) estimates to key methodological decisions embedded in the legacy approach adopted in the U.S. National Weather Service Hydrometeorological Reports No. 51 and No. 52. Although widely used for infrastructure design and risk regulation, fundamental aspects of PMP estimation—such as storm sample size, transposition domain, maximization procedures, and storm rotation—remain poorly constrained and lack formal guidance. Using the Red Rock watershed in Iowa as a case study, and leveraging the 2002–2023 NOAA Analysis of Record for Calibration (AORC) precipitation dataset, we systematically evaluate how each methodological choice, individually and in combination, influences PMP estimates. Our findings demonstrate that PMP is not a fixed physical upper bound but rather a modeling construct shaped heavily by user-defined assumptions. Notably, PMP values derived from modern gridded rainfall datasets can be substantially higher than the legacy estimate used in the original spillway design for Red Rock Dam. Decisions regarding storm sample size, domain extent, climatological window, and particularly storm rotation all contributed to higher PMP estimates. Storm rotation alone—a loosely constrained element in the current PMP practice—can amplify PMP by more than 25%. These results reveal the lack of standardized bounds in current PMP workflows and the need for systematic sensitivity and uncertainty analysis. As PMP estimation shifts toward probabilistic approaches, incorporating physically meaningful storm attributes will be key to developing more transparent, defensible methods for dam safety and climate-resilient infrastructure.

Probable maximum precipitation↗

Antiviral discovery using sparse datasets by integrating experiments, molecular simulations, and machine learning

Computational methods have demonstrated success in identifying virucidal agents, effectively contributing to the discovery of novel virucidal molecules. In this study, we developed a machine learning (ML) model, trained on a small dataset, to predict inhibitors of human enterovirus 71 (EV71), a pathological agent that causes severe disease in children and immunocompromised adults. Despite the dataset’s limitation, comprising of only 36 compounds tested, our ML framework demonstrated significant predictive capability. Notably, experimental validation revealed that five out of the eight compounds predicted by our model from the Chinese cosmetic material list exhibited virucidal activity. The inhibitor effects displayed by the main active compounds were further confirmed by molecular dynamics simulation. This underscores the potential of our AI-driven approach to bypass data constraints in identifying active molecules against viral pathogens.

60 APPLIED LIFE SCIENCES↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

Machine learning pipeline for denoising low signal-to-noise ratio and out-of-distribution transmission electron microscopy datasets

High-resolution transmission electron microscopy (HRTEM) is crucial for observing material’s structural and morphological evolution at Angstrom scales, but the electron beam can alter these processes. Devices such as CMOS-based direct-electron detectors operating in electron-counting mode can be utilized to substantially reduce the electron dosage. However, the resulting images often lead to a low signal-to-noise ratio, which requires frame integration that sacrifices temporal resolution. Several machine learning (ML) models have been recently developed to successfully denoise HRTEM images. Yet, these models are often computationally expensive, and their inference speeds on GPUs are outpaced by the imaging speed of advanced detectors, precluding in situ analysis. Furthermore, the performance of these denoising models on datasets with imaging conditions that deviate from the training datasets has not been evaluated. To mitigate these gaps, we propose a new self-supervised ML denoising pipeline specifically designed for time-series HRTEM images. This pipeline integrates a blind-spot convolution neural network with pre-processing and post-processing steps, including drift correction and low-pass filtering. Results demonstrate that our model outperforms various other ML and non-ML denoising methods in noise reduction and contrast enhancement, leading to improved visual clarity of atomic features. Additionally, the model is drastically faster than U-Net-based ML models and demonstrates excellent out-of-distribution generalization. The model’s computational inference speed is in the order of milliseconds per image, rendering it suitable for application in in-situ HRTEM experiments.

36 MATERIALS SCIENCE↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

Aerial imagery dataset of lost oil wells

Orphaned wells are wells for which the operator is unknown or insolvent. The location of hundreds of thousands of these wells remain unknown in the United States alone. Cost-effective techniques are essential to locate orphaned wells to address environmental problems. In this paper, we present a dataset consisting of 120,948 aerial images of recently documented orphan wells. Each of these 512 × 512 images is paired with segmentation masks that indicate the presence or absence of such well. These images, sourced from the National Agriculture Imagery Program, cover the continental United States with spatial resolutions ranging from 30 centimeters to 1 meter. Additionally, we included negative examples by selecting locations uniformly across the United States. Accompanying metadata includes the IDs and spatial resolution of the original images, which are available for free through the United States Geological Survey, and the pixel coordinates of documented orphaned wells identified in these images. This dataset is intended to support the development of deep-learning models that can help locating undocumented orphan wells from such imagery, thereby blunting the environmental damage they do.

Climate-change mitigation↗

The High-resolution Urban Meteorology for Impacts Dataset (HUMID) daily for the Conterminous United States

Many current gridded surface meteorological datasets are inadequate for quantifying near surface spatiotemporal variability because they do not fully represent the impacts of land surface heterogeneity. Of note, explicit representation of the spatial structure and magnitude of local urban warming are usually lacking. Here we enhance the representation of spatial meteorological variability over urban areas in the conterminous United States (CONUS) by employing the High-Resolution Land Data Assimilation System (HRLDAS), which accounts for the fine-scale impacts of spatiotemporally varying land surfaces on weather. We also synthesize in situ meteorological data including local mesonets to create a 1 km grid spacing model-observation fusion product spanning 1981-2018 over the CONUS. Daily maximum, minimum, and mean values for a variety of temperature estimates, humidity, and surface energy budget terms, among others, are included. This High-resolution Urban Meteorology for Impacts Dataset (HUMID) will be useful for studies examining spatial variability of near surface meteorology and the impacts of urban heat islands across many disciplines including epidemiology, ecology, and climatology.

54 ENVIRONMENTAL SCIENCES↗

Dataset of tensile properties for sub-sized specimens of nuclear structural materials

Mechanical testing with sub-sized specimens plays an important role in the nuclear industry, facilitating tests in confined experimental spaces with lower irradiation levels and accelerating the qualification of new materials. The reduced size of specimens results in different material behavior at the microscale, mesoscale, and macroscale, in comparison to standard-sized specimens, which is referred to as the “specimen size effect.” Although analytical models have been proposed to correlate the properties of sub-sized specimens to standard-sized specimens, these models lack broad applicability across different materials and testing conditions. The objective of this study is to create the first large public dataset of tensile properties for sub-sized specimens used in nuclear structural materials. We performed an extensive literature review of relevant publications and extracted over 1,000 tensile testing records comprising 55 columns including material type and composition, manufacturing information, irradiation conditions, specimen dimensions, and tensile properties. The dataset can serve as a valuable resource to investigate the specimen size effect and develop computational methods to correlate the tensile properties of sub-sized specimens.

36 MATERIALS SCIENCE↗

A co-registered in-situ and ex-situ dataset from wire arc additive manufacturing process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data-centric approach emphasizes leveraging sensor data available throughout the production process to optimize performance. Integration of extensive data analysis provides opportunities for improving precision, reducing waste, and enhancing the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes a comprehensive description of the deposition process, process parameters, welding characteristics and acoustic data collected in-situ, and X-Ray Computed Tomography data of the build.

42 ENGINEERING↗