Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Public Data Set: Impurity Dynamics and Radiative Losses During Local Helicity Injection Startup in the Pegasus-III Spherical Tokamak

This public dataset contains openly-documented, machine readable digital research data corresponding to figures published in C. Rodriguez Sanchez et al., “ Impurity Dynamics and Radiative Losses During Local Helicity Injection Startup in the Pegasus-III Spherical Tokamak,” accepted for publication in Physics of Plasmas .

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Optimal Transport for e/$\pi^0$ Particle Classification in LArTPC Neutrino Experiments

Separation of electron signals from $\pi^0$ backgrounds is crucial for neutrino oscillation measurements and searches for Beyond Standard Model (BSM) physics in current and future Liquid Argon Time Projection Chamber (LArTPC) experiments. e/$\pi^0$ separation has been a reconstruction challenge since both e and $\pi^0$ present as electromagnetic showers, and often only one out of the two showers produced by $\pi^0$ is reconstructed correctly. This research aims to improve the performance of e/$\pi^0$ separation using optimal transport (OT), by leveraging on the topological differences in the showers produced by the two particles. OT is a method which compares two distributions by finding the most efficient way to transform, or “move” from one to the other. This work uses the MicroBooNE open samples public dataset to test the e/$\pi^0$ separation performance of the method on events which incorporate realistic modeling of LArTPC detector response. Reconstructed 3D energy deposits are projected onto a plane perpendicular to the primary shower, allowing OT to better detect the topological differences between the two types of particles without the need to separately reconstruct all the showers in the events. Different distance metrics for OT are tested and preliminary results on e/$\pi^0$ separation are presented.

43 PARTICLE ACCELERATORS↗

One Million Open-source Cislunar Orbits.

The dataset contains one million integrated cislunar orbit trajectories, for a time span of up to six years. The data was generated on LLNL HPC systems and is saved in the form of both HDF5 and CSVs.

Yeager, T↗

Scatter Removal using Black Body Grids and Determining the Minimum Resolvable Hydrogen Concentration at MARS

This dataset was used to develop and validate a scatter-correction pipeline for neutron radiographs acquired at the MARS beamline, based on the methodology of Carminati et al. (2019). Scatter removal was performed using 36 gadolinium black bodies to characterize and remove the spatially varying neutron scatter field that cannot be eliminated by conventional open-beam normalization. The dataset comprises eight samples - four single-crystal nickel and four polycrystalline austenitic 316L stainless steel. Within each material system, samples were pre-charged with hydrogen gas at four charging pressures (1.5, 5, 12.5, and 20 kpsi). Following scatter correction, the pipeline was applied to all radiographs to recover quantitative attenuation maps. After subtracting the known attenuation contribution of the metal matrix, spatial maps of the hydrogen attenuation coefficient were obtained for each sample. This data was then used to construct a calibration curve relating hydrogen attenuation to hydrogen concentration and to determine the minimum resolvable hydrogen concentration at the MARS beamline. The resulting dataset provides a pressure-resolved calibration standard for quantitative hydrogen mapping using neutron radiography.

36 MATERIALS SCIENCE↗

SpaceNet 9—Cross-Sensor Alignment of Optical and SAR Imagery

Precise registration of high-resolution synthetic aperture radar (SAR) and optical imagery is necessary for realizing the full potential and benefits of multimodal image analysis. However, two significant challenges presently exist. First, there is a lack of annotated datasets and benchmarks available for high-resolution SAR–optical image registration. Second, an assessment of efficient and reliable image registration methods that can precisely align these modalities is lacking. Here, we present a holistic description of the SpaceNet 9 Challenge and its results. We present a description of the dataset and baseline algorithm along with the results of the challenge, including a description of the winning algorithms. We release the SpaceNet 9 dataset along with open-sourcing the winning algorithms and baseline. The objective of SpaceNet 9 was to compute a dense displacement map that indicates the shift needed to align pixels in an optical image to the pixels in a SAR image. The challenge launched in April 2025 and was active for approximately two months. The top five solutions reduced image alignment error from approximately 34 m to under 13 m for public and private test data, with the best results obtaining a registration error of only 8.5 and 6.7 m on the public testing and private testing dataset, respectively. Usage of pretrained image matching models, robust outlier rejection with RANSAC, and estimating local displacement were common among the top solutions. The results of this challenge provide insight into high-resolution SAR–optical image registration and offer opportunities for future benchmarking in this domain. The baseline algorithm, winning solutions, and datasets are available at https://spacenet.ai/sn9-challenge/.

benchmark datasets↗

Isolating Unisolated Upsilons with Anomaly Detection in CMS Open Data

We present the first study of anti-isolated Upsilon decays to two muons (ϒ→𝜇⁺⁢𝜇⁻) in proton-proton collisions at the Large Hadron Collider. Using a machine learning (ML)-based anomaly detection strategy, we “rediscover” the ϒ in 13 TeV CMS Open Data from 2016, despite overwhelming anti-isolated backgrounds. We elevate the signal significance to 6.4⁢𝜎 using these methods, starting from 1.6⁢𝜎 using the dimuon mass spectrum alone. Moreover, we demonstrate improved sensitivity from using an ML-based estimate of the multifeature likelihood compared to traditional “cut-and-count” methods. This is the first ever detection of anti-isolated Upsilons, which can be useful in the study of heavy-flavor fragmentation in quantum chromodynamics. Our Letter demonstrates that it is possible and practical to find real signals in experimental collider data using ML-based anomaly detection, and we distill a readily accessible benchmark dataset from the CMS Open Data to facilitate future anomaly detection developments.

machine learning↗

PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE

This dataset is an open-source repository of spectral data measured at Pacific Northwest National Laboratory (PNNL). This database provides quantitative values for the complex index of refraction for seven polycyclic aromatic hydrocarbon (PAH) solids. A list of the chemicals is available in the readme file. These spectra consist of the optical constants, i.e., the real, n(ν), and imaginary, k(ν), refractive indices, over the spectral range from 7,800 to 400 cm-1 (1.28 – 25 μm). The conditions under which the individual data were acquired are described in the associated metadata files, and the user is strongly encouraged to read and understand this information to ensure the data are used appropriately for your application. Recommended Citation for Dataset Jessica M Salcido, Jeremy D. Erickson, Ashley M. Bradley, Russell G. Tonkyn, Timothy J. Johnson and Tanya L. Myers. 2026. PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE. [Data Set] PNNL DataHub. INSERT DOI License Information This work is marked with CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. The authors do request that you appropriately cite the dataset when referencing or using the dataset.

Salcido, Jessica Marie Ortola↗

Bifacial Vertical Testbed and Ground Irradiance Data in Golden, Colorado

This data was collected for Tonita et al., “Vertical bifacial photovoltaic system model validation: study with field data, various orientations, and latitudes,” for validation of optical models for vertically-oriented photovoltaics under high albedo. Ground irradiance data for vertical PV arrays modeling in agrivoltaics is also provided. The dataset is provided for further use or study as open source. For any questions on the dataset, email silvana.ovaitt@nlr.gov.

14 SOLAR ENERGY↗

Towards robust surrogate models: Benchmarking machine learning approaches to expediting phase field simulations of brittle fracture

Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.

Benchmark dataset↗

Drone Flight Data Logs

This dataset represents the open-air tests for the drones when testing different flight scenarios. For some flights we created and tested with a set of onboard sensors. For others we used the native logs for the drones. We recorded relevant conditions for each of the flights to examine environmental issues and weight impacts. We also looked at segmentations of flights to investigate the energy used in each type of flight.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset for "A primer on forest structure measurement with lidar for ecologists"

This repository includes data and code accompanying the case study included in the manuscript "A primer on forest structure measurement with lidar for ecologists" (submitted to Ecosphere). We compiled lidar datasets from multiple platforms in a common area to: 1. Demonstrate how differences in sensor characteristics influence density and resolution of lidar data. 2. Provide open-source, co-located datasets for users to further inspect differences in lidar data. 3. Provide example code to perform basic lidar analysis. This case study is meant to allow readers to get hands-on experience with real-world data from different platforms. This case study is not meant to be a rigorous comparison of derived ecological metrics among all sensors; such comparisons can be found throughout other publications referenced throughout the main manuscript. Code includes basic functions in R commonly used to visualize and manipulate lidar data accessible with a normal laptop computer; more sophisticated algorithms for advanced users are also referenced throughout the main manuscript. Terrestrial laser scanning (TLS), mobile laser scanning (MLS), UAS laser scanning (ULS), airborne laser scanning (ALS), and spaceborne laser scanning (SLS) data were collected within the Smithsonian Environmental Research Center (SERC) forest dynamics plot in Maryland, USA. TLS, MLS, and ALS data were collected within 1 month of the 2021 growing season; ULS data were collected in November 2020 (“leaf-off” data) and July 2022 (“leaf-on” data).

54 ENVIRONMENTAL SCIENCES↗

Cooperative Transmission Expansion Planning Experiment Data and Results

GO WEST is an open-source power grid modeling framework for U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BA) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of U.S. Western Interconnection created by Texas A&M University. Users can try and select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within GO WEST framework. In this sense, TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. TEP minimizes the total cost of the system which comprises the operational cost of satisfying electricity demand (i.e., generation cost), cost of loss of load (i.e., unserved energy), cost of power flow, and cost of new transmission capacity additions (i.e., investment cost). In order to use TEP model, users need to create scenarios with GO WEST framework. In this analysis, outputs from several models are used to create future inputs to GO WEST and TEP models, including GCAM-USA, TELL, CERF and reV. This dataset includes experiment inputs and outputs from three different transmission expansion scenarios (cooperative, intermediate, and individual) for 2019 and 2059. For 2019, a base scenario to illustrate the default (i.e., historical) power grid operations is also included. This study utilizes rcp45hotter_ssp3 scenario from a previous version of GCAM-USA simulations. Sources of the shapefiles in supplementary data are HIFLD Open and U.S. Energy Atlas. Please see the README file for a detailed description of the main and supplementary data.

Capacity Expansion Model↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Hub stability in the calcium calmodulin-dependent protein kinase II

The calcium calmodulin protein kinase II (CaMKII) is a multi-subunit ring assembly with a central hub formed by the association domains. There is evidence for hub polymorphism between and within CaMKII isoforms, but the link between polymorphism and subunit exchange has not been resolved. Here, we present near-atomic resolution cryogenic electron microscopy (cryo-EM) structures revealing that hubs from the α and β isoforms, either standalone or within an β holoenzyme, coexist as 12 and 14 subunit assemblies. Single-molecule fluorescence microscopy of Venus-tagged holoenzymes detects intermediate assemblies and progressive dimer loss due to intrinsic holoenzyme lability, and holoenzyme disassembly into dimers upon mutagenesis of a conserved inter-domain contact. Molecular dynamics (MD) simulations show the flexibility of 4-subunit precursors, extracted in-silico from the β hub polymorphs, encompassing the curvature of both polymorphs. The MD explains how an open hub structure also obtained from the β holoenzyme sample could be created by dimer loss and analysis of its cryo-EM dataset reveals how the gap could open further. An assembly model, considering dimer concentration dependence and strain differences between polymorphs, proposes a mechanism for intrinsic hub lability to fine-tune the stoichiometry of αβ heterooligomers for their dynamic localization within synapses in neurons.

59 BASIC BIOLOGICAL SCIENCES↗

ORBIT-2 Dataset for Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

This dataset release corresponds to the work conducted in ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling, where large-scale AI methods were applied to improve climate and weather resolution. The collection integrates four widely used, publicly available datasets: ERA5, PRISM, DAYMET, and IMERG. To prepare the data for ORBIT-2 model training and evaluation, we applied a preprocessing pipeline that generates paired low-resolution and high-resolution samples, enabling supervised downscaling experiments. The transformation from coarse to fine scales was performed using bilinear regridding, consistent with the procedures described in WeatherBench2, a community benchmark for weather and climate AI models. This dataset supports the development and evaluation of foundation models designed for weather and climate downscaling at exascale. Additional details on methodology and applications can be found in Wang et al., ORBIT-2 (arXiv:2505.04802, 2025).

54 ENVIRONMENTAL SCIENCES↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

High-n Rydberg transition spectroscopy for heavy impurity transport studies in W7-X (invited)

Here, we present a novel spectroscopy approach to investigate impurity transport by analyzing line-radiation following high-n Rydberg transitions. While high-n Rydberg states of impurity ions are unlikely to be populated via impact excitation, they can be accessed by charge exchange (CX) reactions along the neutral beams in high-temperature plasmas. Hence, localized radiation of highly ionized impurities, free of passive contributions, can be observed at multiple wavelengths in the visible range. For the analysis and modeling of the observed Rydberg transitions, a technique for calculating effective emission coefficients is presented that can well reproduce the energy dependence seen in datasets available on the OPEN-ADAS database. By using the rate coefficients and comparing modeling results with the new high-n Rydberg CX measurements, impurity transport coefficients are determined with well-documented 2σ confidence intervals for the first time. This demonstrates that high-n Rydberg spectroscopy provides important constraints on the determination of impurity transport coefficients. By additionally considering Bolometer measurements, which provide constraints on the overall impurity emissivity and, therefore, impurity densities, error bars can be reduced even further.

Instruments & Instrumentation↗