Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ARM Trajectories Data Set Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s ARM Trajectories Data Set (ARMTRAJ) Value-Added Product (VAP) provides trajectory data sets initialized at ARM deployment coordinates and configured using ARM data sets. The four trajectory data sets support aerosol, cloud, and planetary boundary-layer research. Trajectory calculations use the Hybrid Single-Particle Lagrangian Integrated Trajectory (HYSPLIT) model informed by the European Centre for Medium-Range Weather Forecasts (ECMWF) fifth-generation atmospheric reanalysis (ERA5) data set at its highest spatial resolution (~31 km). HYSPLIT also runs at multiple initial starting locations surrounding ARM deployments (in latitude/longitude and/or vertical coordinates), facilitating an ensemble for each sample in the data sets. The ensemble mean and variability reported in ARMTRAJ improve the fidelity and provide uncertainty estimates of trajectory coordinates, thermodynamic properties, and other output fields.

54 ENVIRONMENTAL SCIENCES

ARM Trajectories Data Set Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s ARM Trajectories Data Set (ARMTRAJ) Value-Added Product (VAP) provides trajectory data sets initialized at ARM deployment coordinates and configured using ARM data sets. The six trajectory data sets support aerosol, cloud, planetary boundary layer, and related research (aerosol-cloud interactions, etc.), as well as studies using ARM Aerial Facility (AAF) and tethered balloon system (TBS) measurements. Trajectory calculations use the Hybrid Single-Particle Lagrangian Integrated Trajectory (HYSPLIT) model informed by the European Centre for Medium-Range Weather Forecasts (ECMWF) fifth-generation atmospheric reanalysis (ERA5) data set at its highest spatial resolution (~31 km). HYSPLIT also runs at multiple initial starting locations surrounding ARM deployments (in latitude/longitude and/or vertical coordinates), facilitating an ensemble for each sample in the data sets. The ensemble mean and variability reported in ARMTRAJ improve the fidelity and provide uncertainty estimates of trajectory coordinates, thermodynamic properties, and other output fields.

54 ENVIRONMENTAL SCIENCES

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]

Autogenerating a Domain-Specific Question-Answering Data Set from a Thermoelectric Materials Database to Enable High-Performing BERT Models

We present a method for autogenerating a large domain-specific question-answering (QA) dataset from a thermoelectric materials database. We show that a small language model, BERT, once fine-tuned on this automatically generated dataset of 99,757 QA pairs about thermoelectric materials, affords better performance in the field of thermoelectric materials compared to a BERT model fine-tuned on the generic English-language QA data set, SQuAD-v2. We further show that mixing the two data sets (ours and SQuAD-v2), which have significantly different syntactic and semantic scopes, allows the BERT model to achieve even better performance. The best-performing BERT model fine-tuned on the mixed data set outperforms the models fine-tuned on the other two data sets by scoring an exact match of 67.93% and an F1 score of 72.29% when evaluated on our test data set. This has important implications as it demonstrates the ability to realize high-performing small language models, with modest computational resources, empowered by domain-specific materials data sets which can be generated according to our method.

biological databases

Data Set Analysis to Reduce Uncertainty in Formula Assignments of Ultrahigh Resolution Mass Spectra

Environmental samples contain a vast array of organic compounds with diverse elemental compositions and heteroatom content. Molecular formula assignments of ultrahigh resolution mass spectra (HRMS) hold promise for elucidating the molecular composition of these compounds. However, the need to account for an assortment of heteroatoms increases the uncertainty associated with individual assignments – and ultimately the ecological, biological, and biogeochemical insights gleaned from the assignments. To address this challenge, we introduce a formula assignment strategy that leverages HRMS data sets to improve assignment confidence, filter false assignments, and mitigate bias in assignment routines. The strategy, implemented using CoreMS, first identifies the highest confidence assignment for a recurring ion in a data set by assessing the mass accuracy and isotopologue similarity of all assignments to the ion across the data set. The second component of the strategy examines the consistency of mass errors for an assigned ion throughout a data set and flags formulas with statistically unlikely deviations in mass error. Here, we illustrate the application and utility of the strategy by comparing its results against documented misassignment patterns within a set of oceanographic samples that were measured with 21 T Fourier Transform Ion Cyclotron Resonance Mass Spectrometry. Because the efficacy of our strategy improves with data set size, it is particularly useful for enhancing assignment confidence in large HRMS data sets common in studies of environmental systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure

A Public Data Set of Auto-Generated Geotagged PV Site Equipment, Generated via Deep Learning

In this research, we present a data set over 100 photovoltaic (PV) sites in TX, which have been automatically geotagged via a fully autonomous deep learning (DL) pipeline. Specifically, locations of inverters, tracker/fixed tilt rows, batteries, and substations are labeled algorithmically. To ensure high data quality, all systems have been reviewed manually and any deep learning errors have been corrected. This public data set, as well as the open-sourced pipeline used to generate it, is valuable for site planning, modelling, and insurance purposes. Given time and resources, we hope to extend the data set to additional states/regions in the US.

14 SOLAR ENERGY

Sensitivity of Regional WRF‐Chem Air Quality and Weather Simulations to Biomass‐Burning Emission Data Sets: A Case Study of the Impact of Canadian Wildfire on the US°

This study focuses on the period from June 26 to 29, 2023, when record‐breaking Canadian wildfires severely impacted air quality in the Midwest United States. Using the Weather Research and Forecasting Model with Chemistry (WRF‐Chem) and four biomass‐burning data sets (Fire Inventory from NCAR version 1, Fire Inventory from NCAR version 2.5, Quick Fire Emissions Data set [QFED], and Regional ABI‐VIIRS Emission), we analyzed aerosol transport from Canada to the US and assessed the model's accuracy in predicting PM 2.5 , O 3 , CO and aerosol weather feedback. Model simulations were compared with ground‐based and remote sensing observations as well as field measurements from the Community Research on Climate and Urban Science (CROCUS) project. Our findings show that the movement of a low‐pressure system from the Great Lakes to the Atlantic, combined with the high‐pressure system over the Atlantic, caused the transport of aerosols from Canadian wildfires to the US. Results show WRF‐Chem significantly underestimated key atmospheric components: aerosol optical depth (AOD) by over 50%, PM 2.5 by 65%–90% and peak O 3 concentrations by 50%–55% across four biomass burning data sets. Additionally, CO and NO 2 concentrations were underpredicted. The substantial underestimation of PM 2.5 led to an overestimation of temperature by up to 3.6 °C primarily due to excessive downward shortwave radiation, which resulted from the underestimation of direct aerosol effects and an increase in sensible heat flux. Among the biomass‐burning data sets, QFED produced the most accurate AOD and PM 2.5 predictions due to improved wildfire emission estimates, leading to a 1.0 to 1.5 °C reduction in temperature overestimation during the daytime. These findings underscore the need for improving wildfire emission estimates for trace gases and aerosols to enhance air quality and weather feedback predictions.

WRF-chem model

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY

A Comparison of Three Neodymium Atomic Data Sets for Kilonova Modeling

We examine the impact of input neodymium (Nd) atomic data on the light curves and spectra of kilonovae (KNe), probing the sensitivity of kilonova observables to the atomic physics of this important lanthanide element. We use the SuperNu Monte Carlo radiative transfer code, simulating a simple semianalytic 1D kilonova (KN) with a pure Nd atmosphere, fixing the radiative transfer method while using input atomic data generated by three different codes: the LANL suite of atomic physics codes, HULLAC, and Autostructure. We see that the choice of atomic data significantly shapes the resulting light curves and spectra. Peak bolometric luminosities differ by a ratio of nearly 1.5 between HULLAC/Autostructure and LANL data sets. Moreover, we observe significant near- to mid-IR differences in the structure of the spectra. We specifically attribute these differences to the choice of atomic data for neutral Nd I. Many of the results here have been adapted from a presentation at “Radiative Transfer and Atomic Physics of Kilonovae” in Stockholm, 2023. We additionally present a LANL data set with energies calibrated to available values in the NIST Atomic Spectra Database, and demonstrate that this calibration also significantly affects IR spectral structure at late time. The substantial differences in KN observables that arise from tuning the atomic data of just one lanthanide element highlight the special attention that must be paid to atomic physics uncertainties when modeling KNe, from AT2017gfo to beyond.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

On-the-fly data set combinations with RNTuple

With the expected data volume increase for HL-LHC and the even more complex computing challenges set by future colliders, the need for efficient data storage and processing becomes more pressing. ROOT’s next-generation data format and I/O subsystem, RNTTuple, is designed to address these challenges. RNTTuple already demonstrates a clear improvement in storage and I/O efficiency, as well as overall stability and robustness with respect to its predecessor, TTTree. These improvements provide a solid baseline to introduce novel extensions to common high-energy and nuclear physics (HENP) workflows. Notably, many workflows could benefit from the ability to arbitrarily join and chain data set samples at runtime, which could reduce overall storage requirements and improve application runtime and ergonomics. In this paper, we present the RNTupleProcessor, which enables HENP data set combinations with RNTuple. We will discuss the main design considerations, present the interfaces to support data set combinations and show how they integrate in typical workflows.

de Geus, Florine Willemijn [CERN; Twente U., Ensch

Enhancing Streamflow Reanalysis Across the Conterminous US Leveraging Multiple Gridded Precipitation Data Sets

Streamflow observations, essential for various water resource applications, are often unavailable at critical locations in need. Although different models have been proposed to enhance streamflow predictability at ungauged locations, the challenge extends beyond model fidelity. Differences in meteorologic forcing data sets, precipitation in particular, can significantly affect the accuracy of hydrologic predictions. This challenge intensifies across regions characterized by diverse hydro-climatological and geographical conditions, such as in the conterminous US (CONUS) where a single precipitation product struggles to consistently replicate observed hydrographs, particularly peak flow dynamics. To enhance streamflow predictions, we utilize a VIC-RAPID hydrologic modeling framework driven by multiple commonly used meteorological forcing data sets, such as Daymet, PRISM, ST4, AORC, and their hybrids and create multiple sets of 40-year (1980–2019) hourly, daily, and monthly streamflow reanalysis, Dayflow Version 2, for 2.7 million river reaches across the CONUS. Most forcings lead to skillful streamflow performance, except for ST4 in the mountainous west, where severe radar blockage adversely affects the accuracy. The evaluation using over 6,000 hourly stream gauges shows that hourly AORC and ST4 lead to improved annual peak flow performance over Daymet—driven streamflow (Dayflow V1), particularly in smaller basins, highlighting the value of high temporal resolution forcings in hydrologic predictions. Compared with other benchmark data sets like National Water Model V3.0, AORC-driven VIC-RAPID exhibits improved regional streamflow performance, with comparable peak flow representation. We envision that multi-forcing streamflow reanalysis data can inform regions in need of forcing data enhancement, diagnose hydrologic model performance, and benefit diverse water resource applications.

54 ENVIRONMENTAL SCIENCES

Fracture Intersections under Stress: Laboratory Data and Code [Data set]

The connectivity of natural and induced fractures governs the injection and withdrawal of fluids from subsurface reservoirs. Connectivity depends on intersections that control how fluids mix and move through the entire system. Here, we present data sets from 3D X-ray microscopy measurements of simple fracture networks under stress. 3D printing was used to create prismatic blocks that formed fracture networks composed of 2 orthogonal fractures. The network orientation was either "x" or "+" relative to an applied vertical stress. 3D data sets were collected for normal loads of 25, 100 and 200 Newtons for samples with fracture surfaces with either correlated or uncorrelated asperity distributions. The file contains data from the 12 samples analyzed along with an example code used to extract the intersection geometry. Additional experimental details can be found in the manuscript "Geologic Stress Modulates Fluid Mixing at Fracture Intersections" (10.1038/s43247-026-03525-9)and supplemental information to appear in Communications Earth & Environment in 2026.

02 PETROLEUM

Public Data Set: Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection

This public data set contains openly-documented, machine readable digital research data corresponding to figures published in G.M. Bodner et al., ‘Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection,’ Physics of Plasmas 28, 102504 (2021) and its erratum in Physics of Plasmas 31, 129904 (2024).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Public Data Set: Erratum: “Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection” [Phys. Plasmas 28, 102504 (2021)]

This public data set contains openly-documented, machine readable digital research data corresponding to figures published in G.M. Bodner et al., ‘Erratum: “Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection” [Phys. Plasmas 28, 102504 (2021)],’ Physics of Plasmas 31, 129904 (2024).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE

The System for Classification of Low-Pressure Systems (SyCLoPS): An All-In-One Objective Framework for Large-Scale Data Sets

We propose the first unified objective framework (SyCLoPS) for detecting and classifying all types of low-pressure systems (LPSs) in a given data set. We use the state-of-the-art automated feature tracking software TempestExtremes (TE) to detect and track LPS features globally in ERA5 and compute 16 parameters from commonly found atmospheric variables for classification. A Python classifier is implemented to classify all LPSs at once. The framework assigns 16 different labels (classes) to each LPS data point and designates four different types of high-impact LPS tracks, including tracks of tropical cyclone (TC), monsoonal system, subtropical storm and polar low. The classification process involves disentangling high-altitude and drier LPSs, differentiating tropical and non-tropical LPSs using novel criteria, and optimizing for the detection of the four types of high-impact LPS. A comparison of our labels with those in the International Best Track Archive for Climate Stewardship (IBTrACS) revealed an overall accuracy of 95% in distinguishing between tropical systems, extratropical cyclones, and disturbances. SyCLoPS produces a better TC detection skill compared to the previous algorithms, highlighted by an approximately 6% reduction in the false alarm rate compared to the previous TE algorithm. The vertical cross section composite of the four types of high-impact LPS we detect each shows distinct structural characteristics. Finally, we demonstrate that SyCLoPS is valuable for investigating various aspects of LPSs in climate data, such as the evolution of a single LPS track, patterns of LPS frequencies, and precipitation or wind influence associated with a particular LPS class.

54 ENVIRONMENTAL SCIENCES