Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter↗

QuARC: Development of a Service to Enable FAIR-er Metadata

The ARC Project: The ARC Team located at NASA’s Marshall Space Flight Center conducts quality assessments of metadata records that catalog NASA’s collection of over 9,000 Earth observation data products, stored in a centralized database called the Common Metadata Repository (CMR). The ARC Team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency with the goal of making NASA’s data products more discoverable, accessible, and usable. ARC = Analysis and Review of the CMR

Earth Science Informatics↗

In Interactive, Web-Based Approach to Metadata Authoring

NASA's Global Change Master Directory (GCMD) serves a growing number of users by assisting the scientific community in the discovery of and linkage to Earth science data sets and related services. The GCMD holds over 8000 data set descriptions in Directory Interchange Format (DIF) and 200 data service descriptions in Service Entry Resource Format (SERF), encompassing the disciplines of geology, hydrology, oceanography, meteorology, and ecology. Data descriptions also contain geographic coverage information, thus allowing researchers to discover data pertaining to a particular geographic location, as well as subject of interest. The GCMD strives to be the preeminent data locator for world-wide directory level metadata. In this vein, scientists and data providers must have access to intuitive and efficient metadata authoring tools. Existing GCMD tools are not currently attracting. widespread usage. With usage being the prime indicator of utility, it has become apparent that current tools must be improved. As a result, the GCMD has released a new suite of web-based authoring tools that enable a user to create new data and service entries, as well as modify existing data entries. With these tools, a more interactive approach to metadata authoring is taken, as they feature a visual "checklist" of data/service fields that automatically update when a field is completed. In this way, the user can quickly gauge which of the required and optional fields have not been populated. With the release of these tools, the Earth science community will be further assisted in efficiently creating quality data and services metadata. Keywords: metadata, Earth science, metadata authoring tools

Pollack, Janine↗

An Overview of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI): Supporting FAIR and Open Access to Airborne and Field Data

Since 2019, NASA’s Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) has worked to promote and ensure the discoverability and accessibility of the agency’s non-satellite Earth science observations. A primary component of this effort is the development of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI) and the vetting of key contextual details required to sustain this unique inventory of airborne and field metadata. CASEI provides information on the science objectives motivating data collection, key events/time periods in the observational record aligned with the science objectives, complementary simultaneous observations, programmatic details, and much more. The diverse set of data formats and disciplines served by CASEI have required the implementation of a common data model to organize suborbital observation metadata and efficiently connect appropriate campaigns, platforms, and instruments. The CASEI inventory provides a single entry point for users to search and browse NASA’s airborne and field data archives, regardless of which repository is responsible for their stewardship. This presentation will provide a summary of the motivations for and the development of the CASEI system. Particular attention will be granted to how CASEI facilitates discovery and reuse of these lesser-known NASA data, supporting the Open Science vision and enhancing the return on investments made to collect these unique and varied observations. An up-to-date summary of CASEI inventory content and initial metrics will be provided. Current and future avenues ADMG is pursuing to enhance both CASEI and specific components of suborbital data stewardship at various stages of the data life cycle will also be discussed.

Stephanie M. Wingo↗

Steps Toward Improved Integration, Search, and Analysis of Heterogeneous Data in the Astrobiology Habitable Environments Database

The Astrobiology Habitable Environments Database (AHED) is a new data system being developed as a long-term, open-access repository for astrobiology data. AHED is intended to store user-contributed results from NASA or externally-funded research in astrobiology, and to encourage sharing and synergy within the astrobiology community. However, the interdisciplinary nature of astrobiology presents some specific challenges to data management, integration, and analysis within AHED. In some disciplines (e.g., genomics), open databases thrive because the contributed products are fairly uniform and standardized (e.g., sequence data). In astrobiology, each investigation produces a unique set of data products; this makes it difficult to search across different datasets to find similar data, or to combine results from separate investigations. With AHED, we are taking steps to ensure there is adequate metadata - both at the dataset and record levels - to facilitate search, integration, and analysis. At the dataset level, we are developing a new metadata standard for describing astrobiology datasets, with detailed information about content, funding source, and scientific relevance, along with a set of topical keywords for characterizing datasets. At the record level, we are encouraging users to provide more structured content and finer-grained metadata. In many user-contributed science data repositories, few restrictions are placed on the uploaded data format, and minimal or no record-level metadata is required; thus users are unburdened when it comes to data preparation. The tradeoff is that deep integration and search across datasets is almost impossible without standardized structures and metadata. Although AHED users are free to upload minimally-described datasets, they will be encouraged to use database authoring tools (supplied by the underlying platform - Open Data Repository's Data Publisher) plus a set of customizable astrobiology-specific templates to help structure their data and provide standardized metadata. In reward for their extra effort, AHED will be able to deliver enhanced search, discovery, and analysis capabilities.

astrobiology↗

A User-Focused Renovation of CERES Metadata

Production software and public data products for Clouds and the Earth’s Radiant Energy System (CERES) continue to evolve as the project extends its climate data record. The data management team for CERES is currently undertaking major renovations of both code and data products, the latter of which is, of course, in service of improving user experience. A major mode of CERES’ data product improvement is in renovating products’ metadata. Metadata standards have evolved since CERES began producing its data products in 2000. In its twentieth year, CERES essentially asked the question: how would the project design its data products if it could start all over again? With forthcoming editions, this rebirth will be realized. CERES has redesigned its metadata standards to best position itself for data discoverability. The project has used the latest standards being developed in NASA’s Earth Science Data and Information Systems (ESDIS) Project’s Unified Metadata Model (UMM) documentation; collaborated with the Atmospheric Science Data Center (ASDC) to ensure compliance with Common Metadata Repository compatibility, and continued compliance with Climate and Forecast (CF) Conventions. In doing so, the team created its own, internal document for proper metadata creation and metadata verification software that is deployed prior to all code deliveries. This presentation will discuss this redesign process, as well as needs met and those that are still outstanding in the search for an improved user experience with CERES data products.

Kathleen Dejwakh↗

pyQuARC: Open Source Library for Earth Observation Metadata Quality Assessment

Metadata quality is essential to effective data discovery and has become increasingly vital as more Earth Science data sets become available. The Common Metadata Repository (CMR) hosts metadata describing NASA’s Earth Observation data products, which are archived across 12 Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, conducts metadata quality assessments to ensure that these data products are discoverable, accessible, and usable. To achieve these goals, the ARC team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency. ARC uses a combination of manual and automated methods to assess these three components and identify areas of improvement; the team then collaborates with the DAACs to resolve any findings. To streamline this process, ARC is currently developing a host of scripts, known as pyQuARC, to automate metadata quality assessments as much as possible. pyQuARC is an open source library for Earth Observation Metadata Quality Assessment, and the tool utilizes ARC’s metadata quality assessment framework to make basic validation checks, pinpoint inconsistencies between dataset-level (i.e. collection) and file-level (i.e. granule) metadata, and identify opportunities for more descriptive and robust information. Since pyQuARC is also customizable, other users can make modifications as needed, and future metadata standards can also be implemented. Once pyQuARC is fully developed, it will support multiple schema types to serve the broader EOSDIS metadata community. This presentation will provide an overview of pyQuARC and its process of development while showcasing the tool’s valuable features and uses.

Jenny Wood↗

Optimizing Sample Collection and Accessibility through the Biospecimen and Tissue Sharing Collection (BTSC) Program

The Space Radiation Element (SRE) of the Human Research Program (HRP) is dedicated to establishing a robust biospecimen and tissue sharing collection (BTSC) program that enhances sample collection, tracking, access, distribution, and usability, with the goal of maximizing scientific return. By leveraging biospecimens and tissues from previous experiments, HRP effectively achieves its scientific objectives in characterizing and mitigating the human health impacts of spaceflight while optimizing resource utilization. To further improve the usability and accessibility of the current biospecimen archive, the project aims to expand upon NASA's existing resources and institutional knowledge, ensuring ongoing modernization. To facilitate seamless navigation of the program's workflow, an educational series on the BTSC program is provided to Principal Investigators (PIs). This comprehensive series equips PIs with crucial information on submitting their inventory via the BTSC Metadata Intake Form, ultimately leading to the public availability of their data on NASA's Life Science Portal (NLSP). Covering various aspects such as metadata submission instructions and backend processes for transferring metadata to the Laboratory Information Management System (LIMS), the series incorporates guidance from NASA's Biological Institutional Scientific Collection (NBISC) and Ames Life Sciences Data Archive (ALSDA). The BTSC program represents a significant stride towards enhancing the usability and accessibility of biospecimens for space research. By enabling NASA to deepen its understanding of the health implications of long-term spaceflight, this initiative plays a pivotal role in ensuring the safety and well-being of astronauts.

Shelita Renee Augustus↗

ECHO Status for International Partners

The EOS Clearinghouse (ECHO) is a clearinghouse of spatial and temporal metadata, inclusive of NASA's Distributed Active Archive Center (DAAC) data holdings, that enables the science community to more easily exchange NASA data and information. Currently, ECHO has metadata descriptors for over 55 million individual data granules and 13 million browse images. The majority of ECHO's holdings come directly from data held in the NASA DAACs. The science disciplines and domains represented in ECHO are diverse and include metadata for all of NASA's Science Focus Area data. As middleware for a service-oriented enterprise, ECHO offers access to its capabilities through a set of publicly available Application Program Interfaces (APIs). More information about ECHO is available at http://eos.nasa.gov.echo. The presentation will discuss the status of the ECHO Partners, holdings, and activities, including the transition from the EOS Data Gateway to the Warehouse Inventory Search Tool (WIST)

Weinstein, Beth↗

Enabling Exchange and Adequate Use of Data for Observation Based Atmospheric Research

Systematic long-term field observations have played a vital role in advancing atmospheric research over the past several decades. The use of these observations has expanded from primarily characterizing atmospheric processes and trends to evaluating satellite measurements, assessing models, and improving air quality forecasts. Consequently, the demand for atmospheric chemistry observational data have dramatically increased in terms of scope and coverage of measurements (i.e., parameters/species, spatiotemporal extent). In addition to high quality measurements, certain data reporting standards need to be agreed to ensure the data can be readily exchanged and are sufficiently documented to enable adequate use in different research activities. To this end, WMO has developed and implemented measurement guidelines and community practices for meteorology, climatology, atmospheric and hydrological sciences. In addition, the WMO Expert Team on Metadata Standards manages and evolves the existing metadata standards for the WMO Information System WIS and WMO Integrated Global Observing System WIGOS to support consistent and interoperable data descriptions, ensure relevance to research, and to apply data science principles. This team draws on a wide range of expertise from the research community, including atmospheric measurements, modeling, data management, and data science. The current activities include development of key performance indicators, vocabularies for metadata and the evolution of metadata standards to lower the barrier of application to weather/climate/water/environment data for all communities and the weather enterprise. This presentation intends to promote awareness of ongoing progress and actively solicit community feedback.

Field Observations↗

ESIP Documentation Cluster Session: GCMD Keyword Update

The Global Change Master Directory (GCMD) Keywords are a hierarchical set of controlled Earth Science vocabularies that help ensure Earth science data and services are described in a consistent and comprehensive manner and allow for the precise searching of collection-level metadata and subsequent retrieval of data and services. Initiated over twenty years ago, the GCMD Keywords are periodically analyzed for relevancy and will continue to be refined and expanded in response to user needs. This talk explores the current status of the GCMD keywords, the value and usage that the keywords bring to different tools/agencies as it relates to data discovery, and how the keywords relate to SWEET (Semantic Web for Earth and Environmental Terminology) Ontologies.

data discover↗

Reanalysis of Rat Data from Spacelab Life Sciences 2 (SLS-2) to Reveal Research Gaps in Spaceflight Data

Using and analyzing the legacy data obtained in space life sciences missions has the potential to provide researchers a complete picture of the molecular changes associated with space without further experimentation. This project’s objective is to extract, filter, organize, and analyze all Rattus norvegicus data and metadata obtained from Columbia’s Spacelab Life Sciences 2 (SLS-2, STS-58) mission to explore the ways that we can compile information from model organisms, in our case rats, to create a reliable model to understand biological mechanisms in response to these space flight changes. By reusing rare space legacy data coupled with data analysis techniques, we can combine individual preexisting datasets with current ones to gain new, comprehensive insights about the effects of spaceflight on our bodies. Our methods can also lead to the creation of a standardized pipeline that could be applied to other space life science datasets for analysis. In this review, every biological experiment conducted on rats in the SLS-2 Mission was studied with our pipeline to create a new biological library and model that could be used by scientists from around the world to make novel discoveries and develop new hypotheses from this priceless information without the limitation of the costs of spaceflight experimentation.

rats↗

Development of an Improved Spatial Metadata Simplification Algorithm

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata↗

Cassini/Huygens Program Archive Plan for Science Data

The purpose of this document is to describe the Cassini/Huygens science data archive system which includes policy, roles and responsibilities, description of science and supplementary data products or data sets, metadata, documentation, software, and archive schedule and methods for archive transfer to the NASA Planetary Data System (PDS).

NASA Planetary Science Data System↗

Community Involvement in Enhancing the Global Change Master Directory (GCMD) Controlled Vocabularies (Keywords)

NASA's Global Change Master Directory (GCMD) develops and expands a hierarchical set of controlled vocabularies (keywords) covering the Earth sciences and associated information (data centers, projects, platforms, instruments, etc.). The purpose of the keywords is to describe Earth science data and services in a consistent and comprehensive manner, allowing for the precise searching of metadata and subsequent retrieval of data and services. The keywords are accessible in a standardized SKOSRDFOWL representation and are used as an authoritative taxonomy, as a source for developing ontologies, and to search and access Earth Science data within online metadata catalogues. The keyword development approach involves: (1) receiving community suggestions, (2) triaging community suggestions, (3) evaluating the keywords against a set of criteria coordinated by the NASA ESDIS Standards Office, and (4) publication/notification of the keyword changes. This approach emphasizes community input, which helps ensure a high quality, normalized, and relevant keyword structure that will evolve with users changing needs. The Keyword Community Forum, which promotes a responsive, open, and transparent processes, is an area where users can discuss keyword topics and make suggestions for new keywords. The formalized approach could potentially be used as a model for keyword development.

governance↗

Exploiting Dark Information Resources to Create New Value Added Services to Study Earth Science Phenomena

This paper presents two research applications exploiting unused metadata resources in novel ways to aid data discovery and exploration capabilities. The results based on the experiments are encouraging and each application has the potential to serve as a useful standalone component or service in a data system. There were also some interesting lessons learned while designing the two applications and these are presented next.

Earth Science Informatics↗