Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data publication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Toward a Common Earth Data Publication Framework

Data publication is an essential activity for all data archives. Each of NASA's twelve Distributed Active Archive Centers (DAACs) have established publication workflows which account for the heterogeneous suite of missions, instruments, data providers, and datasets managed within the Earth Observation System Data and Information System (EOSDIS) program. Some aspects of data publication vary across DAACs: workflows range from manual to automatic, terms used to describe publication elements differ, and systems used to publish and manage data vary. Despite these differences, the DAAC data publication processes are generally the same: obtain the data and related information from data providers, describe the data with metadata and documentation, and release the data for access by the user community. In order to improve consistency and reduce the time required to publish data, we have developed a cross-DAAC initiative called the Common Earthdata Publication Framework (Earthdata Pub). Earthdata Pub seeks to: standardize communications and interactions with data providers; identify and standardize common workflows and steps in the data publication process; and design/implement a front-end system with features that include a common web interface, email & status tracking, and common application programming interfaces (APIs) to communicate with various DAAC-specific software components (services and applications) on the back-end. We will present the latest updates on this effort's progress and future plans.

data publication

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)

Collaborative Data Publication Utilizing the Open Data Repository's (ODR) Data Publisher

Introduction: For small communities in diverse fields such as astrobiology, publishing and sharing data can be a difficult challenge. While large, homogenous fields often have repositories and existing data standards, small groups of independent researchers have few options for publishing standards and data that can be utilized within their community. In conjunction with teams at NASA Ames and the University of Arizona, the Open Data Repository's (ODR) Data Publisher has been conducting ongoing pilots to assess the needs of diverse research groups and to develop software to allow them to publish and share their data collaboratively. Objectives: The ODR's Data Publisher aims to provide an easy-to-use and implement software tool that will allow researchers to create and publish database templates and related data. The end product will facilitate both human-readable interfaces (web-based with embedded images, files, and charts) and machine-readable interfaces utilizing semantic standards. Characteristics: The Data Publisher software runs on the standard LAMP (Linux, Apache, MySQL, PHP) stack to provide the widest server base available. The software is based on Symfony (www.symfony.com) which provides a robust framework for creating extensible, object-oriented software in PHP. The software interface consists of a template designer where individual or master database templates can be created. A master database template can be shared by many researchers to provide a common metadata standard that will set a compatibility standard for all derivative databases. Individual researchers can then extend their instance of the template with custom fields, file storage, or visualizations that may be unique to their studies. This allows groups to create compatible databases for data discovery and sharing purposes while still providing the flexibility needed to meet the needs of scientists in rapidly evolving areas of research. Research: As part of this effort, a number of ongoing pilot and test projects are currently in progress. The Astrobiology Habitable Environments Database Working Group is developing a shared database standard using the ODR's Data Publisher and has a number of example databases where astrobiology data are shared. Soon these databases will be integrated via the template-based standard. Work with this group helps determine what data researchers in these diverse fields need to share and archive. Additionally, this pilot helps determine what standards are viable for sharing these types of data from internally developed standards to existing open standards such as the Dublin Core (http://dublincore.org) and Darwin Core (http://rs.twdg.org) metadata standards. Further studies are ongoing with the University of Arizona Department of Geosciences where a number of mineralogy databases are being constructed within the ODR Data Publisher system. Conclusions: Through the ongoing pilots and discussions with individual researchers and small research teams, a definition of the tools desired by these groups is coming into focus. As the software development moves forward, the goal is to meet the publication and collaboration needs of these scientists in an unobtrusive and functional way.

easy to use and implement software tool

Collaborative Data Publication Utilizing the Open Data Repository's Data Publisher

For small communities in multidisciplinary fields such as astrobiology, publishing and sharing data can be challenging. While large, homogenous fields often have repositories and existing data standards, small groups of independent researchers have few options for publishing data that can be utilized within their community. In conjunction with teams at NASA Ames and the University of Arizona, a number of pilot studies are being conducted to assess the needs of these research groups and to guide the software development so that it allows them to publish and share their data collaboratively.

Human-readable interfaces

EARTHDATA PUB: A Data Publication Workflow Solution for NASA’s EOSDIS

Each NASA Distributed Active Archive Center (DAAC) faces the challenge of dealing with an increasingly diverse number of publishable data products from diverse data producers. Data producers, on the other hand, may experience pain points when interacting with the EOSDIS for the first time or when publishing different data at different DAACs. As a result, there has been a growing need to develop a common software framework that serves as a common interface for data producers, rigorously defines the data publication procedure for DAAC staff, facilitates the management of various data publication processes, and tracks the progress of data publication. This software should also account for the different configurations at different DAACs. Currently, two primary data publication workflow and tracking tools exist in operation at the EOSDIS: Semi-Automated ingest System (SAuS) and Data Publication workflow Portal (DAPPeR). However, neither tool is cloud-ready. Automated data processing could be managed by Cumulus, an EOSDIS cloud-based data ingest, archive and management system. However, Cumulus does not support manual tasks or on-premise implementations. We propose to develop the Earthdata Publication Minimum Viable Product (Earthdata Pub MVP) -- a cloud-hosted solution that works with both cloud and on-premise systems and implements the communications and exchange requirements generated by the Earthdata Pub information architecture team.

Rice, Justin L.

Public Data Set: A Magnetic Diagnostic Suite for the Pegasus-III Experiment

This public data set contains openly-documented, machine readable digital research data corresponding to figures published in J.A. Reusch et al., 'A Magnetic Diagnostic Suite for the Pegasus-III Experiment,' Review of Scientific Instruments 95, 093518 (2024). DOI: 10.1063/5.0219341

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Public Data Set: Effects of Injected Current Streams on MHD Equilibrium Reconstruction of Local Helicity Injection Plasmas in a Spherical Tokamak

This public data set contains openly-documented, machine readable digital research data corresponding to figures published in J.D. Weberski et al., 'Effects of Injected Current Streams on MHD Equilibrium Reconstruction of Local Helicity Injection Plasmas in a Spherical Tokamak,' Journal of Fusion Energy 43, 72 (2024).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Public Data Set: Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection

This public data set contains openly-documented, machine readable digital research data corresponding to figures published in G.M. Bodner et al., ‘Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection,’ Physics of Plasmas 28, 102504 (2021) and its erratum in Physics of Plasmas 31, 129904 (2024).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

A Public Data Set of Auto-Generated Geotagged PV Site Equipment, Generated via Deep Learning

In this research, we present a data set over 100 photovoltaic (PV) sites in TX, which have been automatically geotagged via a fully autonomous deep learning (DL) pipeline. Specifically, locations of inverters, tracker/fixed tilt rows, batteries, and substations are labeled algorithmically. To ensure high data quality, all systems have been reviewed manually and any deep learning errors have been corrected. This public data set, as well as the open-sourced pipeline used to generate it, is valuable for site planning, modelling, and insurance purposes. Given time and resources, we hope to extend the data set to additional states/regions in the US.

14 SOLAR ENERGY

Public Data Set: Erratum: “Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection” [Phys. Plasmas 28, 102504 (2021)]

This public data set contains openly-documented, machine readable digital research data corresponding to figures published in G.M. Bodner et al., ‘Erratum: “Initial Characterization of Electron Temperature and Density Profiles in PEGASUS Spherical Tokamak Discharges Driven Solely by Local Helicity Injection” [Phys. Plasmas 28, 102504 (2021)],’ Physics of Plasmas 31, 129904 (2024).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

GEONEX Data Products: Geostationary Satellite Derived Land Surface and Atmospheric Public Data Products

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GEONEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GEONEX pipeline will begin to produce provisional data products to be consumed by external collaborators and the academic community. In order to better inform and introduce the GEONEX products to the science community, the provisional products shall be distributed from the NAS data portal, located at data.nas.nasa.gov, and simple webpage at www.nasa.gov/geonex has been deployed, which describes any algorithms used in deriving the products, user manuals and data file information. We will also update the status of the data processing, on the website and provide links to the latest datasets, and use a geonex mailing list, using lists.nasa.gov.

Wang, Weile

Public Data Set: Impurity Dynamics and Radiative Losses During Local Helicity Injection Startup in the Pegasus-III Spherical Tokamak

This public dataset contains openly-documented, machine readable digital research data corresponding to figures published in C. Rodriguez Sanchez et al., “ Impurity Dynamics and Radiative Losses During Local Helicity Injection Startup in the Pegasus-III Spherical Tokamak,” accepted for publication in Physics of Plasmas .

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin