Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

1235 Preparing for TEMPO: A Review of Planned Metadata, Data Structure, and Distribution by NASA’s Atmospheric Science Data Center

The Atmospheric Science Data Center (ASDC) is in the Science Directorate located at the NASA Langley Research Center (LaRC), in Hampton, Virginia. The ASDC is one of NASA’s Distributed Active Archive Centers (DAAC) and supports over 60 projects and provides access to more than 1,000 archived collections. These datasets were created from satellite measurements, field experiments, and modeled data products. ASDC projects focus on the following Earth science disciplines: Radiation Budget, Clouds, Aerosols, and Tropospheric Composition. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the upcoming Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument.. The instrument will share a ride on a commercial satellite as a hosted payload and will be launched to an orbit about 22,000 miles above Earth's equator. The investigation will, for the first time, use a space-based instrument to make accurate observations of tropospheric pollution concentrations of ozone, nitrogen dioxide, formaldehyde, and aerosols with high resolution and frequency over the U.S, Canada, and Mexico.

Ashlee Autore

1235 Preparing for TEMPO: A Review of Planned Metadata, Data Structure, and Distribution by NASA’s Atmospheric Science Data Center

The Atmospheric Science Data Center (ASDC) is in the Science Directorate located at the NASA Langley Research Center (LaRC), in Hampton, Virginia. The ASDC is one of NASA’s Distributed Active Archive Centers (DAAC) and supports over 60 projects and provides access to more than 1,000 archived collections. These datasets were created from satellite measurements, field experiments, and modeled data products. ASDC projects focus on the following Earth science disciplines: Radiation Budget, Clouds, Aerosols, and Tropospheric Composition. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the upcoming Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument.. The instrument will share a ride on a commercial satellite as a hosted payload and will be launched to an orbit about 22,000 miles above Earth's equator. The investigation will, for the first time, use a space-based instrument to make accurate observations of tropospheric pollution concentrations of ozone, nitrogen dioxide, formaldehyde, and aerosols with high resolution and frequency over the U.S, Canada, and Mexico.

Ashlee Autore

Laying The Foundations for FAIR-ER Science: ISA And LSDA Data Submission Process in NASA’s Evolving Data Management Environment

The Life Sciences Data Archive (LSDA) archives data resulting from research on the effects of spaceflight on humans and the development of countermeasures to mitigate spaceflight hazards. Archivists work with researchers to ensure that unique and high value data products and their metadata are preserved and managed to support current and future research. Currently, LSDA is updating its procedures and data submission requirements in response to the evolving data preservation environment at NASA. LSDA is implementing best practices for research data management through the establishment of clear data submission guidelines, integration of the FAIR (Findability, Accessibility, Interoperability, Reusability) principles, and use of the ISA (Investigation, Study, Assay) research metadata framework for data discoverability and transparency into the data management processes. These changes directly impact LSDA’s requirements for research data submissions. The newly revised Research Data Submission Agreement (RDSA), formerly the Data Submission Agreement (DSA), introduces ISA-compatible metadata collection standards to LSDA’s process. Adherence to LSDA’s data submission guidelines enhances the FAIR-ness of the repository’s collections for future users. This presentation will discuss (1) how submission of research data and associated metadata are impacted by current data management policies, (2) benefits of the adoption of FAIR principles and the ISA metadata framework for retrospective studies utilizing existing LSDA datasets and historic data collections, and (3) the support LSDA will provide to researchers during this transition.

Data submission

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter

QuARC: Development of a Service to Enable FAIR-er Metadata

The ARC Project: The ARC Team located at NASA’s Marshall Space Flight Center conducts quality assessments of metadata records that catalog NASA’s collection of over 9,000 Earth observation data products, stored in a centralized database called the Common Metadata Repository (CMR). The ARC Team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency with the goal of making NASA’s data products more discoverable, accessible, and usable. ARC = Analysis and Review of the CMR

Earth Science Informatics

In Interactive, Web-Based Approach to Metadata Authoring

NASA's Global Change Master Directory (GCMD) serves a growing number of users by assisting the scientific community in the discovery of and linkage to Earth science data sets and related services. The GCMD holds over 8000 data set descriptions in Directory Interchange Format (DIF) and 200 data service descriptions in Service Entry Resource Format (SERF), encompassing the disciplines of geology, hydrology, oceanography, meteorology, and ecology. Data descriptions also contain geographic coverage information, thus allowing researchers to discover data pertaining to a particular geographic location, as well as subject of interest. The GCMD strives to be the preeminent data locator for world-wide directory level metadata. In this vein, scientists and data providers must have access to intuitive and efficient metadata authoring tools. Existing GCMD tools are not currently attracting. widespread usage. With usage being the prime indicator of utility, it has become apparent that current tools must be improved. As a result, the GCMD has released a new suite of web-based authoring tools that enable a user to create new data and service entries, as well as modify existing data entries. With these tools, a more interactive approach to metadata authoring is taken, as they feature a visual "checklist" of data/service fields that automatically update when a field is completed. In this way, the user can quickly gauge which of the required and optional fields have not been populated. With the release of these tools, the Earth science community will be further assisted in efficiently creating quality data and services metadata. Keywords: metadata, Earth science, metadata authoring tools

Pollack, Janine

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)

An Overview of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI): Supporting FAIR and Open Access to Airborne and Field Data

Since 2019, NASA’s Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) has worked to promote and ensure the discoverability and accessibility of the agency’s non-satellite Earth science observations. A primary component of this effort is the development of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI) and the vetting of key contextual details required to sustain this unique inventory of airborne and field metadata. CASEI provides information on the science objectives motivating data collection, key events/time periods in the observational record aligned with the science objectives, complementary simultaneous observations, programmatic details, and much more. The diverse set of data formats and disciplines served by CASEI have required the implementation of a common data model to organize suborbital observation metadata and efficiently connect appropriate campaigns, platforms, and instruments. The CASEI inventory provides a single entry point for users to search and browse NASA’s airborne and field data archives, regardless of which repository is responsible for their stewardship. This presentation will provide a summary of the motivations for and the development of the CASEI system. Particular attention will be granted to how CASEI facilitates discovery and reuse of these lesser-known NASA data, supporting the Open Science vision and enhancing the return on investments made to collect these unique and varied observations. An up-to-date summary of CASEI inventory content and initial metrics will be provided. Current and future avenues ADMG is pursuing to enhance both CASEI and specific components of suborbital data stewardship at various stages of the data life cycle will also be discussed.

Stephanie M. Wingo

Steps Toward Improved Integration, Search, and Analysis of Heterogeneous Data in the Astrobiology Habitable Environments Database

The Astrobiology Habitable Environments Database (AHED) is a new data system being developed as a long-term, open-access repository for astrobiology data. AHED is intended to store user-contributed results from NASA or externally-funded research in astrobiology, and to encourage sharing and synergy within the astrobiology community. However, the interdisciplinary nature of astrobiology presents some specific challenges to data management, integration, and analysis within AHED. In some disciplines (e.g., genomics), open databases thrive because the contributed products are fairly uniform and standardized (e.g., sequence data). In astrobiology, each investigation produces a unique set of data products; this makes it difficult to search across different datasets to find similar data, or to combine results from separate investigations. With AHED, we are taking steps to ensure there is adequate metadata - both at the dataset and record levels - to facilitate search, integration, and analysis. At the dataset level, we are developing a new metadata standard for describing astrobiology datasets, with detailed information about content, funding source, and scientific relevance, along with a set of topical keywords for characterizing datasets. At the record level, we are encouraging users to provide more structured content and finer-grained metadata. In many user-contributed science data repositories, few restrictions are placed on the uploaded data format, and minimal or no record-level metadata is required; thus users are unburdened when it comes to data preparation. The tradeoff is that deep integration and search across datasets is almost impossible without standardized structures and metadata. Although AHED users are free to upload minimally-described datasets, they will be encouraged to use database authoring tools (supplied by the underlying platform - Open Data Repository's Data Publisher) plus a set of customizable astrobiology-specific templates to help structure their data and provide standardized metadata. In reward for their extra effort, AHED will be able to deliver enhanced search, discovery, and analysis capabilities.

astrobiology

A User-Focused Renovation of CERES Metadata

Production software and public data products for Clouds and the Earth’s Radiant Energy System (CERES) continue to evolve as the project extends its climate data record. The data management team for CERES is currently undertaking major renovations of both code and data products, the latter of which is, of course, in service of improving user experience. A major mode of CERES’ data product improvement is in renovating products’ metadata. Metadata standards have evolved since CERES began producing its data products in 2000. In its twentieth year, CERES essentially asked the question: how would the project design its data products if it could start all over again? With forthcoming editions, this rebirth will be realized. CERES has redesigned its metadata standards to best position itself for data discoverability. The project has used the latest standards being developed in NASA’s Earth Science Data and Information Systems (ESDIS) Project’s Unified Metadata Model (UMM) documentation; collaborated with the Atmospheric Science Data Center (ASDC) to ensure compliance with Common Metadata Repository compatibility, and continued compliance with Climate and Forecast (CF) Conventions. In doing so, the team created its own, internal document for proper metadata creation and metadata verification software that is deployed prior to all code deliveries. This presentation will discuss this redesign process, as well as needs met and those that are still outstanding in the search for an improved user experience with CERES data products.

Kathleen Dejwakh

pyQuARC: Open Source Library for Earth Observation Metadata Quality Assessment

Metadata quality is essential to effective data discovery and has become increasingly vital as more Earth Science data sets become available. The Common Metadata Repository (CMR) hosts metadata describing NASA’s Earth Observation data products, which are archived across 12 Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, conducts metadata quality assessments to ensure that these data products are discoverable, accessible, and usable. To achieve these goals, the ARC team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency. ARC uses a combination of manual and automated methods to assess these three components and identify areas of improvement; the team then collaborates with the DAACs to resolve any findings. To streamline this process, ARC is currently developing a host of scripts, known as pyQuARC, to automate metadata quality assessments as much as possible. pyQuARC is an open source library for Earth Observation Metadata Quality Assessment, and the tool utilizes ARC’s metadata quality assessment framework to make basic validation checks, pinpoint inconsistencies between dataset-level (i.e. collection) and file-level (i.e. granule) metadata, and identify opportunities for more descriptive and robust information. Since pyQuARC is also customizable, other users can make modifications as needed, and future metadata standards can also be implemented. Once pyQuARC is fully developed, it will support multiple schema types to serve the broader EOSDIS metadata community. This presentation will provide an overview of pyQuARC and its process of development while showcasing the tool’s valuable features and uses.

Jenny Wood

Optimizing Sample Collection and Accessibility through the Biospecimen and Tissue Sharing Collection (BTSC) Program

The Space Radiation Element (SRE) of the Human Research Program (HRP) is dedicated to establishing a robust biospecimen and tissue sharing collection (BTSC) program that enhances sample collection, tracking, access, distribution, and usability, with the goal of maximizing scientific return. By leveraging biospecimens and tissues from previous experiments, HRP effectively achieves its scientific objectives in characterizing and mitigating the human health impacts of spaceflight while optimizing resource utilization. To further improve the usability and accessibility of the current biospecimen archive, the project aims to expand upon NASA's existing resources and institutional knowledge, ensuring ongoing modernization. To facilitate seamless navigation of the program's workflow, an educational series on the BTSC program is provided to Principal Investigators (PIs). This comprehensive series equips PIs with crucial information on submitting their inventory via the BTSC Metadata Intake Form, ultimately leading to the public availability of their data on NASA's Life Science Portal (NLSP). Covering various aspects such as metadata submission instructions and backend processes for transferring metadata to the Laboratory Information Management System (LIMS), the series incorporates guidance from NASA's Biological Institutional Scientific Collection (NBISC) and Ames Life Sciences Data Archive (ALSDA). The BTSC program represents a significant stride towards enhancing the usability and accessibility of biospecimens for space research. By enabling NASA to deepen its understanding of the health implications of long-term spaceflight, this initiative plays a pivotal role in ensuring the safety and well-being of astronauts.

Shelita Renee Augustus

ECHO Status for International Partners

The EOS Clearinghouse (ECHO) is a clearinghouse of spatial and temporal metadata, inclusive of NASA's Distributed Active Archive Center (DAAC) data holdings, that enables the science community to more easily exchange NASA data and information. Currently, ECHO has metadata descriptors for over 55 million individual data granules and 13 million browse images. The majority of ECHO's holdings come directly from data held in the NASA DAACs. The science disciplines and domains represented in ECHO are diverse and include metadata for all of NASA's Science Focus Area data. As middleware for a service-oriented enterprise, ECHO offers access to its capabilities through a set of publicly available Application Program Interfaces (APIs). More information about ECHO is available at http://eos.nasa.gov.echo. The presentation will discuss the status of the ECHO Partners, holdings, and activities, including the transition from the EOS Data Gateway to the Warehouse Inventory Search Tool (WIST)

Weinstein, Beth

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING