Engineering PapersSearch

SEARCH · Engineering Papers

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Open-Datasets for WEC Simulation

SAND2025-11468O The Open-Datasets for WEC Simulation is a tool that uses datasets and simulation configuration files to conduct OpenFOAM wave energy converter (WEC) simulations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Chartrand, Chris [Sandia National Lab. (SNL-CA), L

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING

Contribution to Open Molecules Dataset (coordcomplexsampling)

The Open Molecules Dataset is a project led by external collaborators, focused on the production of a wide diversity of molecular chemistries. This project will include high-throughput electronic structure calculations on metal coordination complexes, biomolecules, and electrolytes. Solvation effects, conformers, chemical reactivity, and spin/charge sampling will be pursued for elements on the periodic table up to, but not including the actinides. We aim to contribute code related to sampling for transition metal complexes, in particular, and expand software capabilities related to other thrusts desired.

Taylor, Michael G. [Los Alamos National Laboratory

An open retail boundary dataset for South Korea using open data and computer vision technique

Although delineating retail boundaries is important to explore and comprehend the dynamics of the retail sector, it is hard to find studies specifically addressing it in the South Korean context. This study fills this gap by proposing new retail boundaries across South Korea. To achieve this goal, we employed a variety of retailers and building datasets and proposed a unique computer vision-based framework with a deep ensemble voting technique. As a result, we delineated 6,636 distinct retail boundaries that were validated against existing reference retail boundaries. These newly delineated retail boundaries provide valuable insights for researchers, governments, and other relevant stakeholders by enhancing their understanding of retail geography. This dataset can be used as a foundational resource for analyses on topics such as pandemic recovery, retail gentrification, and the resilience of retail spaces in response to e-commerce growth, ultimately contributing to more robust retail sector research in South Korea.

97 MATHEMATICS AND COMPUTING

Poisson Log-Normal Process for Count Data Prediction

Modeling count data is important in physics and other scientific disciplines, where measurements often involve discrete, non-negative quantities such as photon or neutrino detection events. Traditional parametric approaches can be trained to generate integer-count predictions but may struggle with capturing complex, non-linear dependencies often observed in the data. Gaussian process (GP) regression provides a robust non-parametric alternative to modeling continuous data; however, it cannot generate integer outputs. We propose the Poisson Log-Normal (PoLoN) process, a framework that employs GP to model Poisson log-rates. As in GP regression, our approach relies on the correlations between data points captured via GP kernel structure rather than explicit functional parameterizations. We demonstrate that the PoLoN predictive distribution is Poisson-LogNormal and provide an algorithm for optimizing kernel hyperparameters. Furthermore, we adapt the PoLoN approach to the problem of detecting weak localized signals superimposed on a smoothly varying background - a task of considerable interest in many areas of science and engineering. Our framework allows us to predict the strength, location and width of the detected signals. We evaluate PoLoN's performance using both synthetic and real-world datasets, including the open dataset from CERN which was used to detect the Higgs boson at the Large Hadron Collider. Our results indicate that the PoLoN process can be used as a non-parametric alternative for analyzing, predicting, and extracting signals from integer-valued data.

Saha, Anushka [Rutgers U., Piscataway]

An Open Benchmark of One Million High-Fidelity Cislunar Trajectories

Cislunar space spans from geosynchronous altitudes to beyond the Moon and will underpin future exploration, science, and security operations. We describe and release an open dataset of one million numerically propagated cislunar trajectories generated with the open-source Space Situational Awareness Python package (SSAPy). The model includes high-degree Earth/Moon gravity, solar gravity, and Earth/Sun radiation pressure; other planetary gravities are omitted by design for computational efficiency. Initial conditions uniformly sample commonly used osculating-element ranges, and each trajectory is propagated for up to six years under a single, fixed start epoch. The dataset is intended as a reusable benchmark for method development (e.g., space domain awareness, navigation, and machine-learning pipelines), a reference library for statistical studies of orbit families, and a starting point for community-driven extensions (e.g., alternative epochs). We report empirically observed stability trends (e.g., a band near ~5 GEO and persistence of some co-orbital classes including L4/L5 librators) as dataset descriptors rather than new dynamical results. The chief contribution is the scale, fidelity, organization (CSV/HDF5 with full state time series and metadata), and open availability, which together lower the barrier to comparative and data-driven studies in the cislunar regime.

79 ASTRONOMY AND ASTROPHYSICS

Visual and Inertial Datasets for an eVTOL Aircraft Approach and Landing Scenario

A National Aeronautics and Space Administration (NASA) project developing computer vision algorithms for autonomous flight is producing real-world datasets with cameras mounted on aircraft. In related domains, such as autonomous driving, open datasets are key to innovation and advancement in computer vision and autonomous perception for future Advanced Air Mobility (AAM) operations. Few vision datasets, however, are publicly available in the aviation context. This paper introduces preliminary datasets containing several examples of approach and landing scenarios. The platform aircraft include a multirotor small unmanned aerial system (sUAS) and a crewed helicopter as surrogates for future electric vertical take-off and landing (eVTOL) aircraft. The dataset provides video imagery with associated inertial navigation system-global positioning system (INS-GPS) position and attitude estimates and other sensors. Surveyed locations of the visual features of the landing area are included. This dataset is the first to be released in an ongoing effort to collect and share large, diverse datasets relevant to autonomous aviation; community critique that can inform and improve future flight campaigns is welcome.

Nelson Brown

Discovery of correlated electron molecular orbital materials using graph representations

Correlated electron molecular orbital (CEMO) materials host emergent electronic states built from molecular orbitals localized over clusters of transition metal ions yet have historically been discovered sporadically and generally been treated as isolated case studies. Here we establish CEMO materials as a systematically discoverable class and introduce a graph-based framework to identify, classify, and organize transition-metal cluster motifs in inorganic solids. Starting from crystal structures in the Materials Project, we construct transition metal connectivity graphs, extract cluster motifs using a bond-cutting algorithm, and determine cluster point groups, effective cluster sublattice dimensionality, and translational symmetry. Applying this approach in a high-throughput screen of 34,548 compounds yields 5,306 cluster-containing materials, including 2,627 stable or metastable compounds with isolated clusters and 984 materials featuring mixed-metal clusters. The resulting dataset reveals symmetry and element dependent trends in cluster formation. By integrating cluster classification with flat band lattice topology and battery-relevant information, we provide further relevant information to multiple scientific communities. The accompanying open dataset, Cluster Finder software, and interactive web platform enable systematic exploration of cluster driven electronic phenomena and establish a general pathway for discovering correlated quantum materials and functional materials with cluster-based or extended metal-metal bonding in inorganic solids.

Akhond, Md. Rajbanul [Department of Chemistry, 800

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields, in part due to a culture of open data sharing and reuse. AI/ML methodology is well-suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Inexperienced researchers can produce models that perform poorly outside of the training dataset. Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Casaletto

IMAGE Software Suite

The IMAGE Mission is generating a truely unique set of magnetospheric measurement through a first-of-its-kind complement of remote, global observations. These data are being distributed in the Universal Data Format (UDF), which consists of data, calibration, and documentation. This is an open dataset, available to all by request to the National Space Science Data Center (NSSDC) at NASA Goddard Space Flight Center. Browse data, which consists of summary observations, is also available through the NSSDC in the Common Data Format (CDF) and graphic representations of the browse data. Access to the browse data can be achieved through the NSSDC CDAWeb services or by use of NSSDC provided software tools. This presentation documents the software tools, being provided by the IMAGE team, for use in viewing and analyzing the UDF telemetry data. Like the IMAGE data, these tools are openly available. What these tools can do, how they can be obtained, and how they are expected to evolve will be discussed.

Gallagher, Dennis L.

Ten questions on building stock modeling to inform energy efficiency and sustainability

To enhance economic competitiveness and ensure energy efficiency, resilience, and security, cities and governments are adopting technologies and strategies to improve their existing building stocks. This approach aims to reduce energy use, improve energy affordability, and ensure a reliable power supply while safeguarding occupants during extreme weather events that may disrupt energy services. The effectiveness of these solutions will depend on building stock characteristics, use patterns, weather conditions, evolving technologies and their markets, and a city’s socio-economic conditions. This paper presents ten questions and answers that highlight the most important issues regarding the use of building stock modeling as a powerful tool to provide insights for informing stakeholders’ actions and decision-making on energy efficiency, costs reduction, and resilience of buildings in cities. Building stock modeling should build upon the fit-for-purpose framework, balancing the use case accuracy requirements, level of complexity, and needed resources (expertise, compute). The advancements in Artificial Intelligence (AI), the increasingly available open dataset of building stock in cities, and the more affordable powerful computing will accelerate the adoption of building stock modeling across scales by researchers and practitioners to inform decision making on sustainability and efficiency.

AI

Repository of HydroSMADE: Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites

This repository presents HydroSMADE—Hydropower Site-level Monthly Availability Data Ensemble, a new open dataset that provides monthly hydropower availability for 1,593 existing and 124,333 potential sites worldwide over the period 1950–2100. The dataset is generated by using a global hydrologic model (Xanthos) with explicit representation of hydropower operation. Specifically, HydroSMADE distinguishes between storage and diversion sites, applies optimized operating rules, and incorporates site-specific characteristics such as generation capacity, maximum turbine flow, and reservoir storage. Driven by bias-corrected meteorological inputs, the data is provided for 30 alternative future scenarios. The scenarios consist of the full factorial combination of three standard CMIP6 atmospheric forcing pathways (SSP1-2.6, SSP3-7.0, and SSP5-8.5) and ten CMIP6 General Circulation Models (GCMs): GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, EC-Earth3, CanESM5, MIROC6, CNRM-ESM2-1, UKESM1-0-LL, and CNRM-CM6-1. The repository contains a total of 122 files: a text file (readme.txt) containing a brief description of the included data, a CSV file containing site attributes, and the remaining 120 files (in CSV) containing site-level monthly hydropower availability. Example Jupyter Notebooks to explore the HydroSMADE dataset are available on GitHub at https://github.com/kamal0013/HydroSMADE More details on the methods and technical validation of HydroSMADE are available in the following paper by the same authors: Chowdhury, A. K., Abeshu, G. W., Zhao, M., Wild, T. B., Hassan, N., Ying, Z., Kim, G. J., Matthew, B., Jonathan, L., & Li, H.-Y. (Submitted). Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites.

Existing and Potential Sites

Public-Private Partnerships to Enable Discovery, Access, and Use of NASA’s Open Earth Science Datasets

Knowledge transfer between public research institutes and private entities is an essential component of the open science movement. While both public and private institutions are making research advances in technologies, organizational boundaries can hinder knowledge transfer. Productive public-private collaboration frameworks are needed to advance research further.

Elizabeth Fancher

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Open Source Synergy: Developing and Validating PMU Data Analysis Techniques Using Open Source Tools and Datasets

This paper presents an exploration into the development and validation of data analysis approaches for Phasor Measurement Units (PMUs) using open-source datasets and tools. Various methods for event detection, event classification, frequency response, and oscillation analysis were tested. We leverage the capabilities of Archive Walker (AW), the Frequency Response Analysis Tool (FRAT), and the Oscillation Baselining and Analysis Tool (OBAT), all open-source tools, for efficient processing and analysis of synchrophasor data. The open-source Transmission Signature Library (TSL) dataset was employed as a dataset for a comprehensive evaluation to assess the performance and reliability of the proposed methods.

PMU, event analysis, oscillation, Frequency Respon