Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data stewardship”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An Overview of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI): Supporting FAIR and Open Access to Airborne and Field Data

Since 2019, NASA’s Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) has worked to promote and ensure the discoverability and accessibility of the agency’s non-satellite Earth science observations. A primary component of this effort is the development of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI) and the vetting of key contextual details required to sustain this unique inventory of airborne and field metadata. CASEI provides information on the science objectives motivating data collection, key events/time periods in the observational record aligned with the science objectives, complementary simultaneous observations, programmatic details, and much more. The diverse set of data formats and disciplines served by CASEI have required the implementation of a common data model to organize suborbital observation metadata and efficiently connect appropriate campaigns, platforms, and instruments. The CASEI inventory provides a single entry point for users to search and browse NASA’s airborne and field data archives, regardless of which repository is responsible for their stewardship. This presentation will provide a summary of the motivations for and the development of the CASEI system. Particular attention will be granted to how CASEI facilitates discovery and reuse of these lesser-known NASA data, supporting the Open Science vision and enhancing the return on investments made to collect these unique and varied observations. An up-to-date summary of CASEI inventory content and initial metrics will be provided. Current and future avenues ADMG is pursuing to enhance both CASEI and specific components of suborbital data stewardship at various stages of the data life cycle will also be discussed.

Stephanie M. Wingo↗

Why We Do What We Do: Data Reuse, Open Access, and Privacy in Data Management at the Life Sciences Data Archive

As custodian of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and how the archive’s evolving data management practices support FAIR-ness; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of of data analysis and aggregation tools.

Data↗

Why We Do What We Do: Data Reuse, Open Access and Privacy in Data Management at the Life Sciences Data Archive

As custodians of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and the archive’s evolving data management practices; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of data analysis and aggregation tools.

data management↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.

Life Sciences data↗

Hanford Site Mule Deer Monitoring Report for Fiscal Years 2024 and 2026

The U.S. Department of Energy, Hanford Field Office (HFO) conducts ecological monitoring on the Hanford Site to collect and track data needed to ensure compliance with environmental laws, regulations, and policies governing Department of Energy activities. The vision for the HFOmanaged portion of the Hanford Site, hereby referred to as Central Hanford, focuses not only on the cleanup of nuclear facilities and waste sites but on the protection and restoration of the Hanford Site lands. As the HFO moves toward accomplishing this vision, understanding of the ecological resources present and the need for conservation and/or protection of those resources will be critical for making informed decisions for responsible site stewardship. Ecological monitoring data provides baseline information about the plants, animals, and habitats under HFO stewardship at Central Hanford required for decision-making under the National Environmental Policy Act of 1969 (NEPA) and Comprehensive Environmental Response, Compensation, and Liability Act of 1980.

54 ENVIRONMENTAL SCIENCES↗

Radiance Data Products at the GES DAAC

The Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) has been archiving and distributing Radiance data, and serving science and application users of these data, for over 10 years now. The user-focused stewardship of the Radiance data from the AIRS, AVHRR, MODIS, SeaWiFS, SORCE, TOMS, TOVS, TRMM, and UARS instruments exemplifies the GES DAAC tradition and experience. Radiance data include raw radiance counts, onboard calibration data, geolocation products, radiometric calibrated and geolocated-calibrated radiance/reflectance. The number of science products archived at the GES DAAC is steadily increasing, as a result of more sophisticated sensors and new science algorithms. Thus, the main challenge for the GES DAAC is to guide users through the variety of Radiance data sets, provide tools to visualize and reduce the volume of the data, and provide uninterrupted access to the data. This presentation will describe the effort at the GES DAAC to build a bridge between multi-sensor data and the effective scientific use of the data, with an emphasis on the heritage of the science products. The intent is to inform users of the existence of this large collection of Radiance data; suggest starting points for cross-platform science projects and data mining activities; provide data services and tools information; and to give expert help in the science data formats and applications.

Savtchenko, A.↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Approach to Managing MeaSURES Data at the GSFC Earth Science Data and Information Services Center (GES DISC)

A major need stated by the NASA Earth science research strategy is to develop long-term, consistent, and calibrated data and products that are valid across multiple missions and satellite sensors. (NASA Solicitation for Making Earth System data records for Use in Research Environments (MEaSUREs) 2006-2010) Selected projects create long term records of a given parameter, called Earth Science Data Records (ESDRs), based on mature algorithms that bring together continuous multi-sensor data. ESDRs, associated algorithms, vetted by the appropriate community, are archived at a NASA affiliated data center for archive, stewardship, and distribution. See http://measures-projects.gsfc.nasa.gov/ for more details. This presentation describes the NASA GSFC Earth Science Data and Information Services Center (GES DISC) approach to managing the MEaSUREs ESDR datasets assigned to GES DISC. (Energy/water cycle related and atmospheric composition ESDRs) GES DISC will utilize its experience to integrate existing and proven reusable data management components to accommodate the new ESDRs. Components include a data archive system (S4PA), a data discovery and access system (Mirador), and various web services for data access. In addition, if determined to be useful to the user community, the Giovanni data exploration tool will be made available to ESDRs. The GES DISC data integration methodology to be used for the MEaSUREs datasets is presented. The goals of this presentation are to share an approach to ESDR integration, and initiate discussions amongst the data centers, data managers and data providers for the purpose of gaining efficiencies in data management for MEaSUREs projects.

Vollmer, Bruce↗

An Overview of NASA’s Airborne and Field Data Resource Center

A key recommendation from NASA’s 2022 Airborne and Field Data Workshop called for the development of a virtual Resource Center for all stakeholders across the data lifecycle of airborne and field Earth observations. The agency’s Earth Science Data and Information System (ESDIS) Project and Airborne Data Management Group (ADMG) have worked in concert to establish the newly launched NASA Airborne and Field Data Resource Center (AFDRC) to provide a single entry point for a wide assortment of information on and the effective, responsible stewardship of non-satellite observational data. The AFDRC compiles access to many existing resources, but does so in a newly organized way that integrates availability to increase efficiency and holistic understanding while simplifying users’ experience. Initially launched in fall of 2023, the AFDRC is a NASA Earthdata domain website that clarifies several previously disparate resources and provides newly updated information, including: Learning Resources: Educational resources to broaden understanding of the role airborne and field observations play in advancing understanding of our planet and NASA’s role in collecting and archiving these data. Support for data users to Find and Access Data: Advanced contextual browse/search capabilities that efficiently link researchers to data products suitable for their science objectives - this includes linking to NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI). Working with Data: Tools specific to individual types of suborbital Earth Science data and their (inter-)disciplinary communities to provide access as well as guidance for their application. Stewardship Responsibilities: Resources for data producers with to lessen requirement burdens at the time of data transfer, and information for data stewards with consistent, authoritative guidance on best practices and agency- and/or community- specific archival procedures. This presentation will give an overview of the motivation for NASA’s AFDRC, approach for the design and content, iterative community-driven improvements, promote the use of the AFDRC, and solicit additional feedback from airborne and field data user communities.

Sara Lubkin↗

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena↗

Materials Data Science Ontology(MDS-Onto): Unifying Domain Knowledge in Materials and Applied Data Science

Ontologies have gained popularity in the scientific community as a way to standardize terminologies in organizations’ data. Although certain cohorts have created frameworks with rules and guidelines on creating ontologies, there exist significant variations in how Materials Science ontologies are currently developed. We seek to provide guidance in the form of a unified automated framework for developing interoperable and modular ontologies for Materials Data Science that simplifies the ontology terms matching by establishing a semantic bridge up to the Basic Formal Ontology(BFO). This framework provides key recommendations on how ontologies should be positioned within the semantic web, what knowledge representation language is recommended, and where ontologies should be published online to boost their findability and interoperability. Two fundamental components of the MDS-Onto framework are the bilingual package called FAIRmaterials for ontology creation and FAIRLinked, for FAIR data creation. To showcase the practical capabilities of FAIRmaterials, we present two exemplar domain ontologies of MDS-Onto: Synchrotron X-Ray Diffraction and Photovoltaics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

FAIRmaterials: Ontology Tools with Data FAIRification in Development

The bilingual FAIRmaterials package simplifies the creation and visualization of materials and data science ontologies. FAIRmaterials, available in the Python and R languages, addresses the complexities associated with traditional ontology editors based on manual user input such as Protege with an intuitive workflow and easy-to-use templates, making it accessible to users both experienced and inexperienced with ontologies. The FAIRmaterials package is its ability to programatically convert simple and structured CSV inputs into rich, well-defined ontologies. This capability is designed to support the findability, accessibility, interoperability, and reusability (FAIR) of research data and serve as a tool in the process of data FAIRification. Its additional features, such as automated ontology merging, static visualizations, and comprehensive documentation for outputs extend its utility, making it a valuable tool for any researcher engaged in knowledge management.

Bradley, Alexander Harding [Case Western Reserve U↗

FAIRLinked: Data FAIRification Tools for Materials Data Science

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of data from CSV into JSONLDs and vice versa in a way that captures the semantics of the data using MDS-Onto. Lastly, QBWorkflow is a serialization and deserialization workflow that incorporates RDF Data Cube vocabulary, useful for working with multidimensional datasets. By offering these packages, FAIRLinked lowers the barrier of creating FAIR, machine-actionable data for researchers in the materials science community.

FAIR↗

Analysis and Review of NASA Earth Science Metadata: How Automation Plays a Role

The Analysis and Review of the Common Metadata Repository (CMR ARC) Team reviews all EOSDIS metadata. The team’s objective is to achieve consistency, correctness, and completeness for all metadata records in the CMR, as well as improve the discoverability of NASA's Earth Science data within the CMR framework. This work is currently being completed at Marshall Space Flight Center. CMR makes a single discovery point possible for NASA's Earth Science data users. The CMR team, in collaboration with three other core metadata teams, contributes to the stewardship of NASA's Earth Science data through a process of continual curation and the ongoing development of the Unified Metadata Model (UMM). A key tool now used in the curation process, referred to as the NASA CMR Dashboard, is an online curation dashboard developed in collaboration with software development company, Element 84. This tool facilitates the review of Earth Science metadata records and subsequent stakeholder collaboration on the resolution of identified issues. A key capability of the new tool is a suite of automated compliance checks written in Python 3.6 that verify the integrity of various metadata elements across multiple standards.

Staton, Patrick↗

Managing and Servicing Physical Oceanographic Data at a NASA Distributed Active Archive Center

The NASA Earth Science Data Information Systems Project funds and operates 12 Distributed Active Archive Center(s) (DAAC) throughout the United States. Of these 12 centers, the Physical Oceanography DAAC (PO.DAAC) is committed to providing long term archival, distribution and stewardship for NASA physical oceanographic data, primarily derived from space-born satellite systems, but also including a growing set of recent and future in situ observations from the SPURS-1 and SPURS-2 campaigns. Notable NASA missions supported include: Seasat, TOPEX/Poseidon, NSCAT, QuikSCAT, ISS-RapidScat, Jason-1, Jason-2/OSTM, GRACE, Aquarius, GHRSST, and MODIS. The following interagency and international missions are also supported by PO.DAAC: AVHRR, Coriolis, DMSP, MetOp-A, MetOp-B, Oceansat-2. The PO.DAAC currently holds 525 datasets in public distribution, spanning the following observational parameters: sea surface temperature, sea surface salinity, ocean color, ocean surface currents, ocean surface wind speed, ocean surface wind direction, sea surface height, significant wave height, ocean water mass/thickness, and sea ice age. A hundred of these datasets are available in near-real-time. Datasets are distributed through a variety of open-source access protocols including FTP, OPeNDAP, and THREDDS. FTP will soon be phased out in favor of a recently introduced HTTPS PO.DAAC Drive interface that supports WebDAV and interoperable machine-to-machine communication. OPeNDAP supports remote data/metadata query, subset, and download. THREDDS provides the features of OPeNDAP with the additional feature of temporal aggregation. PO.DAAC also offers proprietary tools and services to further enhance the data discovery, visualization and analysis experience, including but not limited to: State of the Ocean, Web Services (data/metadata discovery and extraction), HiTIDE Level-2 subsetter, Live Access Server (LAS), Webification (w10nsci), and Rich Site Summary (RSS) Datacasting. To assist with provenance of datasets, PO.DAAC has implemented DOIs for the data it distributes so that they can be properly cited. There is a user forum and helpdesk that contains data recipes and via which users can get guidance. In summary, this presentation aims to provide a general overview of PO.DAAC’s web portal and data holdings along with a set of illustrative examples leading prospective data users into the practical utility of its tools and services.

Moroni, David F.↗

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez↗

A Statistical Model of Tropical Cyclone Tracks in the Western North Pacific with ENSO-Dependent Cyclogenesis

A new statistical model for western North Pacific Ocean tropical cyclone genesis and tracks is developed and applied to estimate regionally resolved tropical cyclone landfall rates along the coasts of the Asian mainland, Japan, and the Philippines. The model is constructed on International Best Track Archive for Climate Stewardship (IBTrACS) 1945-2007 historical data for the western North Pacific. The model is evaluated in several ways, including comparing the stochastic spread in simulated landfall rates with historic landfall rates. Although certain biases have been detected, overall the model performs well on the diagnostic tests, for example, reproducing well the geographic distribution of landfall rates. Western North Pacific cyclogenesis is influenced by El Nino-Southern Oscillation (ENSO). This dependence is incorporated in the model s genesis component to project the ENSO-genesis dependence onto landfall rates. There is a pronounced shift southeastward in cyclogenesis and a small but significant reduction in basinwide annual counts with increasing ENSO index value. On almost all regions of coast, landfall rates are significantly higher in a negative ENSO state (La Nina).

Yonekura, Emmi↗