Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Enabling Exchange and Adequate Use of Data for Observation Based Atmospheric Research

Systematic long-term field observations have played a vital role in advancing atmospheric research over the past several decades. The use of these observations has expanded from primarily characterizing atmospheric processes and trends to evaluating satellite measurements, assessing models, and improving air quality forecasts. Consequently, the demand for atmospheric chemistry observational data have dramatically increased in terms of scope and coverage of measurements (i.e., parameters/species, spatiotemporal extent). In addition to high quality measurements, certain data reporting standards need to be agreed to ensure the data can be readily exchanged and are sufficiently documented to enable adequate use in different research activities. To this end, WMO has developed and implemented measurement guidelines and community practices for meteorology, climatology, atmospheric and hydrological sciences. In addition, the WMO Expert Team on Metadata Standards manages and evolves the existing metadata standards for the WMO Information System WIS and WMO Integrated Global Observing System WIGOS to support consistent and interoperable data descriptions, ensure relevance to research, and to apply data science principles. This team draws on a wide range of expertise from the research community, including atmospheric measurements, modeling, data management, and data science. The current activities include development of key performance indicators, vocabularies for metadata and the evolution of metadata standards to lower the barrier of application to weather/climate/water/environment data for all communities and the weather enterprise. This presentation intends to promote awareness of ongoing progress and actively solicit community feedback.

Field Observations

ESIP Documentation Cluster Session: GCMD Keyword Update

The Global Change Master Directory (GCMD) Keywords are a hierarchical set of controlled Earth Science vocabularies that help ensure Earth science data and services are described in a consistent and comprehensive manner and allow for the precise searching of collection-level metadata and subsequent retrieval of data and services. Initiated over twenty years ago, the GCMD Keywords are periodically analyzed for relevancy and will continue to be refined and expanded in response to user needs. This talk explores the current status of the GCMD keywords, the value and usage that the keywords bring to different tools/agencies as it relates to data discovery, and how the keywords relate to SWEET (Semantic Web for Earth and Environmental Terminology) Ontologies.

data discover

Reanalysis of Rat Data from Spacelab Life Sciences 2 (SLS-2) to Reveal Research Gaps in Spaceflight Data

Using and analyzing the legacy data obtained in space life sciences missions has the potential to provide researchers a complete picture of the molecular changes associated with space without further experimentation. This project’s objective is to extract, filter, organize, and analyze all Rattus norvegicus data and metadata obtained from Columbia’s Spacelab Life Sciences 2 (SLS-2, STS-58) mission to explore the ways that we can compile information from model organisms, in our case rats, to create a reliable model to understand biological mechanisms in response to these space flight changes. By reusing rare space legacy data coupled with data analysis techniques, we can combine individual preexisting datasets with current ones to gain new, comprehensive insights about the effects of spaceflight on our bodies. Our methods can also lead to the creation of a standardized pipeline that could be applied to other space life science datasets for analysis. In this review, every biological experiment conducted on rats in the SLS-2 Mission was studied with our pipeline to create a new biological library and model that could be used by scientists from around the world to make novel discoveries and develop new hypotheses from this priceless information without the limitation of the costs of spaceflight experimentation.

rats

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF

Development of an Improved Spatial Metadata Simplification Algorithm

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING

Cassini/Huygens Program Archive Plan for Science Data

The purpose of this document is to describe the Cassini/Huygens science data archive system which includes policy, roles and responsibilities, description of science and supplementary data products or data sets, metadata, documentation, software, and archive schedule and methods for archive transfer to the NASA Planetary Data System (PDS).

NASA Planetary Science Data System

Community Involvement in Enhancing the Global Change Master Directory (GCMD) Controlled Vocabularies (Keywords)

NASA's Global Change Master Directory (GCMD) develops and expands a hierarchical set of controlled vocabularies (keywords) covering the Earth sciences and associated information (data centers, projects, platforms, instruments, etc.). The purpose of the keywords is to describe Earth science data and services in a consistent and comprehensive manner, allowing for the precise searching of metadata and subsequent retrieval of data and services. The keywords are accessible in a standardized SKOSRDFOWL representation and are used as an authoritative taxonomy, as a source for developing ontologies, and to search and access Earth Science data within online metadata catalogues. The keyword development approach involves: (1) receiving community suggestions, (2) triaging community suggestions, (3) evaluating the keywords against a set of criteria coordinated by the NASA ESDIS Standards Office, and (4) publication/notification of the keyword changes. This approach emphasizes community input, which helps ensure a high quality, normalized, and relevant keyword structure that will evolve with users changing needs. The Keyword Community Forum, which promotes a responsive, open, and transparent processes, is an area where users can discuss keyword topics and make suggestions for new keywords. The formalized approach could potentially be used as a model for keyword development.

governance

Exploiting Dark Information Resources to Create New Value Added Services to Study Earth Science Phenomena

This paper presents two research applications exploiting unused metadata resources in novel ways to aid data discovery and exploration capabilities. The results based on the experiments are encouraging and each application has the potential to serve as a useful standalone component or service in a data system. There were also some interesting lessons learned while designing the two applications and these are presented next.

Earth Science Informatics

Finding Atmospheric Composition (AC) Metadata

The Atmospheric Composition Portal (ACP) is an aggregator and curator of information related to remotely sensed atmospheric composition data and analysis. It uses existing tools and technologies and, where needed, enhances those capabilities to provide interoperable access, tools, and contextual guidance for scientists and value-adding organizations using remotely sensed atmospheric composition data. The initial focus is on Essential Climate Variables identified by the Global Climate Observing System CH4, CO, CO2, NO2, O3, SO2 and aerosols. This poster addresses our efforts in building the ACP Data Table, an interface to help discover and understand remotely sensed data that are related to atmospheric composition science and applications. We harvested GCMD, CWIC, GEOSS metadata catalogs using machine to machine technologies - OpenSearch, Web Services. We also manually investigated the plethora of CEOS data providers portals and other catalogs where that data might be aggregated. This poster is our experience of the excellence, variety, and challenges we encountered.Conclusions:1.The significant benefits that the major catalogs provide are their machine to machine tools like OpenSearch and Web Services rather than any GUI usability improvements due to the large amount of data in their catalog.2.There is a trend at the large catalogs towards simulating small data provider portals through advanced services. 3.Populating metadata catalogs using ISO19115 is too complex for users to do in a consistent way, difficult to parse visually or with XML libraries, and too complex for Java XML binders like CASTOR.4.The ability to search for Ids first and then for data (GCMD and ECHO) is better for machine to machine operations rather than the timeouts experienced when returning the entire metadata entry at once. 5.Metadata harvest and export activities between the major catalogs has led to a significant amount of duplication. (This is currently being addressed) 6.Most (if not all) Earth science atmospheric composition data providers store a reference to their data at GCMD.

metadata search

Use of Spatial Metadata Simplification for TEMPO

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata

NASA's Earth Observing Data and Information System - Near-Term Challenges

NASA's Earth Observing System Data and Information System (EOSDIS) has been a central component of the NASA Earth observation program since the 1990's. EOSDIS manages data covering a wide range of Earth science disciplines including cryosphere, land cover change, polar processes, field campaigns, ocean surface, digital elevation, atmosphere dynamics and composition, and inter-disciplinary research, and many others. One of the key components of EOSDIS is a set of twelve discipline-based Distributed Active Archive Centers (DAACs) distributed across the United States. Managed by NASA's Earth Science Data and Information System (ESDIS) Project at Goddard Space Flight Center, these DAACs serve over 3 million users globally. The ESDIS Project provides the infrastructure support for EOSDIS, which includes other components such as the Science Investigator-led Processing systems (SIPS), common metadata and metrics management systems, specialized network systems, standards management, and centralized support for use of commercial cloud capabilities. Given the long-term requirements, and the rapid pace of information technology and changing expectations of the user community, EOSDIS has evolved continually over the past three decades. However, many challenges remain. Challenges addressed in this paper include: growing volume and variety, achieving consistency across a diverse set of data producers, managing information about a large number of datasets, migration to a cloud computing environment, optimizing data discovery and access, incorporating user feedback from a diverse community, keeping metadata updated as data collections grow and age, and ensuring that all the content needed for understanding datasets by future users is identified and preserved.

Remote Sensing

Operability on the Europa Clipper Mission: Challenges and Opportunities

Flight and ground system operability has been a focus area on the Europa Clipper Project since early in its formulation phase. This has given the operations team the opportunity to influence the design, with a goal of increasing overall system operability. This paper presents example operability challenges, opportunities, and solutions arising from the pre-Critical Design Review (CDR) system design. The integrated wing assembly design directly couples a scientific instrument (the REASON sounding radar) to the spacecraft’s power source (solar array wing panels). Impacts to mission operations of this design include: increased slew durations; solar array pointing constraints during inner cruise, Europa flybys, and orbit trim maneuvers; and stray light intrusions into the stellar reference units’ keep out zones. The use of CCSDS File Delivery Protocol (CFDP) Class-2 for reliable downlink of the large volume of Europa Clipper science data is described, along with nominal and off-nominal use cases. The effort to improve post-launch spacecraft visibility by adding a third low-gain antenna to the spacecraft is detailed. The design of the bulk data store has necessitated the implementation of accountable data products (ADPs), accountability identifiers (AIDs), and metadata packets to provide end-to-end science data accountability. To streamline and automate the flight rules generation and checking process, a first order and temporal logic-based solution of expressing flight rules without ambiguity, and whose programmatic implementation can be automated, is proposed. The focus on operability has had a positive influence on Europa Clipper design decisions, although cost, schedule, budget, heritage, and other technical concerns have many times outweighed operability concerns. However, experience to date demonstrates that this approach to operability results in more thorough, balanced consideration of the effect of early design trades and decisions on the operations phase of a mission than seen in many previous missions, and provides operations development insight into prioritizing work to go.

Signorelli, Joel

Operability on the Europa Clipper Mission: Challenges and Opportunities

Flight and ground system operability has been a focus area on the Europa Clipper Project since early in its formulation phase. This has given the operations team the opportunity to influence the design, with a goal of increasing overall system operability. This paper presents example operability challenges, opportunities, and solutions arising from the Critical Design Review (CDR) system design. The integrated wing assembly design directly couples a scientific instrument (the REASON sounding radar) to the spacecraft’s power source (solar array wing panels). Impacts to mission operations of this design include: increased slew durations; solar array pointing constraints during inner cruise, Europa flybys, and orbit trim maneuvers; and stray light intrusions into the stellar reference units’ keep out zones. The use of CCSDS File Delivery Protocol (CFDP) Class-2 for reliable downlink of the large volume of Europa Clipper science data is described, along with nominal and off-nominal use cases. The effort to improve post-launch spacecraft visibility by adding a third low-gain antenna to the spacecraft is detailed. The design of the bulk data store has necessitated the implementation of accountable data products (ADPs), accountability identifiers (AIDs), and metadata packets to provide end-to-end science data accountability. To streamline and automate the flight rules generation and checking process, a first order and temporal logic-based solution of expressing flight rules without ambiguity, and whose programmatic implementation can be automated, is proposed. The focus on operability has had a positive influence on Europa Clipper design decisions, although cost, schedule, budget, heritage, and other technical concerns have many times outweighed operability concerns. However, experience to date demonstrates that this approach to operability results in more thorough, balanced consideration of the effect of early design trades and decisions on the operations phase of a mission than seen in many previous missions, and provides operations development insight into prioritizing work to go.

Kumar, Meghana

Complications of Metadata Curation for NASA Airborne and Field Campaigns, Platforms, and Instruments

The Airborne Data Management Group (ADMG) curates metadata that describe NASA's airborne and field campaigns, platforms and instruments. This activity is vital to building a useful inventory of sub-orbital Earth science data that improves data discovery and access. During the curation process, many metadata issues were identified that required improvement to campaign and data product metadata. In some cases, locating the needed metadata to add to the inventory was a simple process. For other cases, the information was hard to find. In addition, identifying accurate investigation instrument details to add to the inventory was especially complicated because of the variety of definitions used in the Earth science community for the same concepts. One example of this is the concept of instruments' spatial and temporal resolution. The spatial resolution is one of the more difficult elements to curate given the variations in meaning across various disciplines. Clarified definitions are needed to enable consistency of information across campaigns and instruments. In this presentation, we introduce results from a survey of scientists from various fields in which we asked for definitions of spatial and temporal resolution. Our survey results highlight the importance of creating more universally acceptable definitions for certain metadata elements. By curating sub-orbital field campaign and instrument metadata, ADMG is enabling more efficient discovery and access to NASA observations by allowing science data users to search for certain clearly defined criteria and metadata values.

Ashlyn Shirey

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana

ICARTT File Format Enhancements: Supporting FAIRness and Data Discovery of Suborbital Campaign Data

Suborbital campaigns aim to accomplish a wide variety of goals and can include a variety of platforms, instruments, and parameters measured. In 2004, the ICARTT (International Consortium for Atmospheric Research on Transport and Transformation) standards were developed to fulfill data management needs for the ICARTT campaign. The ICARTT file format is text-based and composed of a header with important data description information and the data section. Built on the NASA Ames and GTE data formats, the ICARTT format was created to facilitate data exchange and promote collaborations among the science teams for achieving the ICARTT campaign goals. Due to its success and adaptation for use in many other field campaigns, the ICARTT file format became a NASA standard in 2010 and was amended in January 2017. These changes provided many enhancements, including the requirement for variable standard names. Primarily designed for airborne field studies, ICARTT has been further utilized for ground-based studies. NASA has made a commitment to build an inclusive open science community over the next decade. Open-source science strives to make publicly funded scientific research transparent, inclusive, accessible, and reproducible. The ICARTT format can host metadata that is critical for proper use of the data, particularly for in-situ measurements, and can enhance data discovery and accessibility. However, the required fields are often free text, meaning that the information is human readable, but not machine interpretable. Furthermore, the amount and type of information provided can vary significantly between principal investigators and campaigns. To support FAIR principles and interoperability, enhancements to the ICARTT standards are recommended. Possible recommendations include potential use of controlled and consistent vocabulary for variable standard name and certain common metadata elements; standardizing timestamps for easier data comparisons and analysis; and providing guidance on variable measurement units and how they are reported. Enhancing ICARTT metadata can further streamline the process to make suborbital data more readily available to the data user and improve variable-level metadata. Providing more variable-level metadata can enhance data searching and discovery, supporting NASA’s Open-Source Science Initiative (OSSI).

Megan Buzanowicz

NASA'S Earth Science Data Stewardship Activities

NASA has been collecting Earth observation data for over 50 years using instruments on board satellites, aircraft and ground-based systems. With the inception of the Earth Observing System (EOS) Program in 1990, NASA established the Earth Science Data and Information System (ESDIS) Project and initiated development of the Earth Observing System Data and Information System (EOSDIS). A set of Distributed Active Archive Centers (DAACs) was established at locations based on science discipline expertise. Today, EOSDIS consists of 12 DAACs and 12 Science Investigator-led Processing Systems (SIPS), processing data from the EOS missions, as well as the Suomi National Polar Orbiting Partnership mission, and other satellite and airborne missions. The DAACs archive and distribute the vast majority of data from NASA’s Earth science missions, with data holdings exceeding 12 petabytes The data held by EOSDIS are available to all users consistent with NASA’s free and open data policy, which has been in effect since 1990. The EOSDIS archives consist of raw instrument data counts (level 0 data), as well as higher level standard products (e.g., geophysical parameters, products mapped to standard spatio-temporal grids, results of Earth system models using multi-instrument observations, and long time series of Earth System Data Records resulting from multiple satellite observations of a given type of phenomenon). EOSDIS data stewardship responsibilities include ensuring that the data and information content are reliable, of high quality, easily accessible, and usable for as long as they are considered to be of value.

metadata