Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data stewardship”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Cumulus: NASA Archives in the Cloud

NASA's Earth Observing System Data and Information System (EOSDIS) houses nearly 30PBs of critical Earth Science data and with upcoming missions is expected to balloon to between 200PBs-300PBs over the next seven years. The magnitude of data collected makes it infeasible to download data and process it locally, forcing us to re-think how we store and work with earth science data. NASA has looked to the cloud to address this, building its open source Cumulus software to manage the ingest of diverse data in a wide variety of formats into the cloud and provide services to manage and access the data. In this talk, we will describe how Cumulus provides common features needed to manage a cloud archive in the realm of ingest, data stewardship, and cost-controlled distribution of science data to users and services.

EOSDIS

NASA's Earth Observing Data and Information System (EOSDIS)

NASA's Earth Observing System Data and Information System (EOSDIS) has been a central component of the NASA Earth observation program since the 1990's. The data collected by NASA represent a significant public investment in research. Consequently, NASA developed a free, open and non-discriminatory policy consistent with existing international policies to maximize access to data. EOSDIS manages data covering a wide range of Earth science disciplines including cryosphere, land cover change, polar processes, field campaigns, ocean surface, geodesy, atmosphere dynamics and composition, and inter-disciplinary research, and many others. This presentation will discuss data stewardship and archive activities.

Data Systems

Introduction to the JPSS-2 Advanced Technology Microwave Sounder (ATMS) Government Calibration Data Book (GCDB)

The third Advanced Technology Microwave Sounder (ATMS) is an instrument onboard the Joint Polar Satellite System (JPSS), JPSS-2 (renamed NOAA-21 in orbit) mission. This report is to introduce the JPSS-2 Government Calibration Data Book (J2 GCDB) for ATMS, SN 304. This J2 GCDB document contains key information generated during the calibration testing campaign that is driving parameters for radiometric performance. This document also contains supporting data that augments the calibration results. The values in this document are utilized by ATMS’s calibration packet which is, in turn, an integral component in the interpretation of science data. The calibration data in this report was collected from tests such as shelf-level testing, antenna testing, instrument thermal vacuum (TVAC) testing; satellite TVAC testing; and JPSS-2 post-launch tests. JPSS-2 was launched on November 10, 2022. In the subsequent years, the Government will release an ATMS GCDB for each JPSS mission. We expect that all public users can download these ATMS GCDBs from the NOAA operational Integrated Calibration and Validation System (ICVS) website, see more discussions below. The goal of this GCDB is to demonstrate how to characterize ATMS measurements using JPSS-2 ATMS on-orbit operational data and to provide relevant explanations. This document serves as a primary public domain reference for calibrating operational ATMS Raw Data Records (RDR) science data, as used in the current operational Interface Data Processing Segment (IDPS) system. This same RDR science data is distributed through direct broadcast (DB) to DB users for use in their ground processing systems. This J2 GCDB provides the results of the ATMS system radiometric calibration, the antenna flat reflector emissivity [1], the antenna pattern measurements, the antenna pattern corrected brightness temperature [2], the brightness temperature of the lunar disk [3], Lunar Intrusion (LI) correction algorithm [4], receiver spectral parameters, and mechanical alignment on-orbit pointing results, and the striping effect appeared significantly in S-NPP on-orbit radiance data when the data are compared to the Radiative Transfer Model (RTM) simulation in numerical weather prediction (NWP) system [5]. It also provides the parameters required for conversion of telemetry counts to engineering units, for radiometric calibration, and for antenna beam geo-location. Moreover, it provides JPSS-2 ATMS Spectral Response Functions data, some additional information related to ATMS on-orbit performance, on-orbit lunar intrusion correction parameters and Earth contamination bias, and on how to derive ATMS RDR, antenna Temperature Data Records (TDR), and Sensor Data Records (SDR). Furthermore, an introduction of NOAA operational Integrated Calibration and Validation System (ICVS) website and services is added in this J2 GCDB. This ICVS hosts a long-term monitoring system which allows to visualization and comparison of data from JPSS missions, NOAA legacy Polar Operational Environmental Satellites (POES), and Geostationary Operational Environmental Satellites (GOES). From NOAA Comprehensive Large Array-data Stewardship System (CLASS), the public users can download all JPSS ATMS data products for all JPSS missions.

Microwave Sounder

AIRS Mission Support from GES DISC

This talk will describe the support and distribution of AIRS (Atmospheric Infra Red Sounding) data products that are archived and distributed from the Goddard Earth Sciences Data and Information Services Center. Along with data stewardship, an important mission of GES DISC is to enhance the usability of data and broaden the user base. We will provide a brief summary of the current online archive and distribution metrics for the AIRS v5 and v6 products. We will also describe collaborative data sets and services (e.g., visualization and potential science applications) and solicit feedback for potential future services.

version

GES DISC Greenhouse Gas Data Sets and Associated Services

NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC) archives and distributes rich collections of data on atmospheric greenhouse gases from multiple missions. Hosted data include those from the Atmospheric Infrared Sounder (AIRS) mission (which has observed CO2, CH4, ozone, and water vapor since 2002); legacy water vapor and ozone retrievals from TIROS Operational Vertical Sounder (TOVS); and Upper Atmosphere Research Satellite (UARS) going back to the early 1980s. GES DISC also archives and supports data from seven projects of the Making Earth System Data Records for Use in Research Environments (MEaSUREs) program that have ozone and water vapor records. Greenhouse gases data from the A-Train satellite constellation is also available: (1) Aura-Ozone Monitoring Instrument (OMI) and Microwave Limb Sounder (MLS) ozone, nitrous oxide, and water vapor since 2004; (2) Greenhouse Gases Observing Satellite (GOSAT) CO2 observations since 2009 from the Atmospheric CO2 Observations from Space (ACOS) task; and (3) Orbiting Carbon Observatory-2 (OCO-2) CO2 data since 2014. The most recent related data set that the GES DISC archives is methane flux for North America, as part of NASAs Carbon Monitoring System (CMS) project. This dataset contains estimates of methane emission in North America based on an inversion of the GEOS-Chem chemical transport model constrained by GOSAT observations (Turner et al., 2015). Along with data stewardship, an important focus area of the GES DISC is to enhance the usability of its data and broaden its user base. Users have unrestricted access to a new user-friendly search interface, which includes many services such as variable subsetting, format conversion, quality screening, and quick browse. The majority of the GES DISC data sets are also accessible through Open-source Project for a Network Data Access Protocol (OPeNDAP) and Web Coverage Service (WCS). The latter two services provide more options for specialized subsetting, format conversion, and image viewing. Additional data exploration, data preview, and preliminary analysis capabilities are available via NASA Giovanni, which obviates the need forusers to download the data (Acker and Leptoukh, 2007). Giovanni provides a bridge between the data and science and has been very successful in extending GES DISC data to educational users and to users with limited resources.

data ordering

Science Data Preservation: Implementation and Why It Is Important

Remote Sensing data generation by NASA to study Earth s geophysical processes was initiated in 1960 with the launch of the first Television Infrared Observation Satellite Program (TIROS), to develop a meteorological satellite information system. What would be deemed as a primitive data set by today s standards, early Earth science missions were the foundation upon which today s remote sensing instruments have built their scientific success, and tomorrow s instruments will yield science not yet imagined. NASA Scientific Data Stewardship requirements have been documented to ensure the long term preservation and usability of remote sensing science data. In recent years, the Federation of Earth Science Information Partners and NASA s Earth Science Data System Working Groups have organized committees that specifically examine standards, processes, and ontologies that can best be employed for the preservation of remote sensing data, supporting documentation, and data provenance information. This presentation describes the activities, issues, and implementations, guided by the NASA Earth Science Data Preservation Content Specification (423-SPEC-001), for preserving instrument characteristics, and data processing and science information generated for 20 Earth science instruments, spanning 40 years of geophysical measurements, at the NASA s Goddard Earth Sciences Data and Information Services Center (GES DISC). In addition, unanticipated preservation/implementation questions and issues in the implementation process are presented.

Kempler, Steven J.

NASA GES DISC support of CO2 Data from OCO-2, ACOS, and AIRS

NASA Goddard Earth Sciences Data and Information Services Centers (GES DISC) is the data center assigned to archive and distribute current AIRS, ACOS data and data from the upcoming OCO-2 mission. The GES DISC archives and supports data containing information on CO2 as well as other atmospheric composition, atmospheric dynamics, modeling and precipitation. Along with the data stewardship, an important mission of GES DISC is to facilitate access to and enhance the usability of data as well as to broaden the user base. GES DISC strives to promote the awareness of science content and novelty of the data by working with Science Team members and releasing news articles as appropriate. Analysis of events that are of interest to the general public, and that help in understanding the goals of NASA Earth Observing missions, have been among most popular practices.Users have unrestricted access to a user-friendly search interface, Mirador, that allows temporal, spatial, keyword and event searches, as well as an ontology-driven drill down. Variable subsetting, format conversion, quality screening, and quick browse, are among the services available in Mirador. The majority of the GES DISC data are also accessible through OPeNDAP (Open-source Project for a Network Data Access Protocol) and WMS (Web Map Service). These services add more options for specialized subsetting, format conversion, image viewing and contributing to data interoperability.

Wei, Jennifer C

Technologies and Methods Used at the Laboratory for Atmospheric and Space Physics (LASP) to Serve Solar Irradiance Data

The Laboratory for Atmospheric and Space Physics (LASP) at the University of Colorado in Boulder, USA operates the Solar Radiation and Climate Experiment (SORCE) NASA mission, as well as several other NASA spacecraft and instruments. Dozens of Solar Irradiance data sets are produced, managed, and disseminated to the science community. Data are made freely available to the scientific immediately after they are produced using a variety of data access interfaces, including the LASP Interactive Solar Irradiance Datacenter (LISIRD), which provides centralized access to a variety of solar irradiance data sets using both interactive and scriptable/programmatic methods. This poster highlights the key technological elements used for the NASA SORCE mission ground system to produce, manage, and disseminate data to the scientific community and facilitate long-term data stewardship. The poster presentation will convey designs, technological elements, practices and procedures, and software management processes used for SORCE and their relationship to data quality and data management standards, interoperability, NASA data policy, and community expectations.

Pankratz, Chris

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) is a facility data management application developed for the NASA Ames arc jet facilities. The current decentralized data management practices limit statistical tracking, synchronization between video/time series, search capability, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management

Big-data Efficient Automated Science Transfer (BEAST): an open-source software architecture for arc jet data management, modeling, and automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management

GHRC: NASAs Hazardous Weather Distributed Active Archive Center

The Global Hydrology Resource Center (GHRC; ghrc.nsstc.nasa.gov) is one of NASA's twelve Distributed Active Archive Centers responsible for providing access to NASA's Earth science data to users worldwide. Each of NASA's twelve DAACs focuses on a specific science discipline within Earth science, provides data stewardship services and supports its research community's needs. Established in 1991 as the Marshall Space Flight Center DAAC and renamed GHRC in 1997, the data center's original mission focused on the global hydrologic cycle. However, over the years, data holdings, tools and expertise of GHRC have gradually shifted. In 2014, a User Working Group (UWG) was established to review GHRC capabilities and provide recommendations to make GHRC more responsive to the research community's evolving needs. The UWG recommended an update to the GHRC mission, as well as a strategic plan to move in the new direction. After a careful and detailed analysis of GHRC's capabilities, research community needs and the existing data landscape, a new mission statement for GHRC has been crafted: to provide a comprehensive active archive of both data and knowledge augmentation services with a focus on hazardous weather, its governing dynamical and physical processes, and associated applications. Within this broad mandate, GHRC will focus on lightning, tropical cyclones and storm-induced hazards through integrated collections of satellite, airborne, and in-situ data sets. The new mission was adopted at the recent 2015 UWG meeting. GHRC will retain its current name until such time as it has built substantial data holdings aligned with the new mission.

Data Archive

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks associated with deep space crewed missions (cis-Lunar, Mars transit/surface) require development of health countermeasures, novel ecosystem support, risk modeling, and fundamental space biological knowledge discovery. Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from space biological and health studies are needed for reuse by scientists to address these tasks. The data as well as space-relevant biospecimens are being made more findable, accessible, interoperable, and reusable through NASA’s Open Science Data Repository (OSDR). This new OSDR umbrella grouping includes NASA GeneLab, the NASA Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection. The OSDR system design appropriately handles metadata and processed-tabular results from ALSDA studies collected from space experiments. But raw and processed ALSDA bioimage and video datasets require an expansion of OSDR’s data architecture to handle ingestion, curation, and egress. The academic-industry bioimaging field saw a scientific renaissance in the past several years through leveraging open-source software, international collaborations, machine learning, and other open science/programming approaches. As crewed missions and more biological experiments are on the deep space horizon, OSDR is embracing data stewardship through listening to feedback from subject matter experts and designing an expanded architecture which is appropriate for NASA’s goals to enable analysis and reuse of bioimaging and video data for the public science community.Discovery Through Image and Video Data Sharing

space biology

Cultivating an Emergent Earth Observation Analytics Ecosystem in the Cloud

A diverse set of data analytics systems for Earth Observations are sprouting up in the Earth Science community, with a wealth of processing algorithms and analysis methods. There is a similar wealth of data resources available via myriad data providers and clearinghouses, including large institutional systems like the Earth Observing System Data and Information System, Comprehensive Large Scale Array-data Stewardship System, and Federated Earth Observation Missions gateway. With Earth system science driving a need to work with more datasets together, and the community developing more analysis tools (some of them dataset-specific), how can we develop analysis workflows that incorporate far-flung datasets and leverage analysis resources from multiple organizations? Cloud computing points the way toward a solution in two different respects. Firstly, the access to and abstraction of virtually unlimited storage and computing power provides an environment that enables more straightforward means of pulling datasets and analysis resources together. Just as importantly, however, cloud computing serves as an example of an "ecosystem" of interoperating services, since the essence of cloud computing is the presentation of all resources as a service, from hardware to infrastructure to platform to software. This enables the combination of off-the-shelf, diverse services to construct entire systems that emerge out of an equally diverse community of architects and developers. This approach can be similarly applied to the data and analysis resources in the Earth Observation community. By exposing these resources via well understood services, and consuming resources in the same way, different organizations can construct bespoke analysis workflows and systems for their own purposes. The key leap the community needs to make is to develop analysis systems in components that interact with other components via services. The result would be a rich ecosystem of analytics components that can be combined to analyze datasets at scale and in conjunction with other datasets from other sources.

chaos

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF

ESnet Data and AI Workshop Report

In February 2025, the DOE user facility Energy Sciences Network (ESnet) held a three-day Data and AI Workshop in Berkeley, California. The objective of the workshop was to identify challenges within ESnet that could be addressed through data-driven methods, to help define ESnet’s data-analysis requirements, and to shape its AI strategy, guiding data-stewardship efforts and the direction of AI research and AIOps exploration for ESnet7, the next iteration of ESnet’s network. This report summarizes the multi-faceted discussions and findings and presents a set of recommendations for next steps.

97 MATHEMATICS AND COMPUTING

Task 51 - Cloud-Optimized Format Study

The cloud infrastructure provides a number of capabilities that can dramatically improve access and use of Earth Observation data. However, in many cases, data may need to be reorganized and/or reformatted in order to make them tractable to support cloud-native analysis/access patterns. The purpose of this study is to examine the pros and cons of different formats for storing data on the cloud. The evaluation will focus on both enabling high-performance data access and usage as well as to meet the existing scientific data stewardship needs of EOSDIS.

Durbin, Chris