Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data

NASA's Global Change Master Directory: Discover and Access Earth Science Data Sets, Related Data Services, and Climate Diagnostics

NASA's Global Change Master Directory provides the scientific community with the ability to discover, access, and use Earth science data, data-related services, and climate diagnostics worldwide. The GCMD offers descriptions of Earth science data sets using the Directory Interchange Format (DIF) metadata standard; Earth science related data services are described using the Service Entry Resource Format (SERF); and climate visualizations are described using the Climate Diagnostic (CD) standard. The DIF, SERF and CD standards each capture data attributes used to determine whether a data set, service, or climate visualization is relevant to a user's needs. Metadata fields include: title, summary, science keywords, service keywords, data center, data set citation, personnel, instrument, platform, quality, related URL, temporal and spatial coverage, data resolution and distribution information. In addition, nine valuable sets of controlled vocabularies have been developed to assist users in normalizing the search for data descriptions. An update to the GCMD's search functionality is planned to further capitalize on the controlled vocabularies during database queries. By implementing a dynamic keyword "tree", users will have the ability to search for data sets by combining keywords in new ways. This will allow users to conduct more relevant and efficient database searches to support the free exchange and re-use of Earth science data. http://gcmd.nasa.gov/

Aleman, Alicia

Collaborative Metadata Curation in Support of NASA Earth Science Data Stewardship

Growing collection of NASA Earth science data is archived and distributed by EOSDIS’s 12 Distributed Active Archive Centers (DAACs). Each collection and granule is described by a metadata record housed in the Common Metadata Repository (CMR). Multiple metadata standards are in use, and core elements of each are mapped to and from a common model – the Unified Metadata Model (UMM). Work done by the Analysis and Review of CMR (ARC) Team.

data stewardship

Evolution in Metadata Quality: Common Metadata Repository's Role in NASA Curation Efforts

Metadata Quality is one of the chief drivers of discovery and use of NASA EOSDIS (Earth Observing System Data and Information System) data. Issues with metadata such as lack of completeness, inconsistency, and use of legacy terms directly hinder data use. As the central metadata repository for NASA Earth Science data, the Common Metadata Repository (CMR) has a responsibility to its users to ensure the quality of CMR search results. This poster covers how we use humanizers, a technique for dealing with the symptoms of metadata issues, as well as our plans for future metadata validation enhancements. The CMR currently indexes 35K collections and 300M granules.

metadata quality

Controlled Vocabularies Boost International Participation and Normalization of Searches

The Global Change Master Directory's (GCMD) science staff set out to document Earth science data and provide a mechanism for it's discovery in fulfillment of a commitment to NASA's Earth Science progam and to the Committee on Earth Observation Satellites' (CEOS) International Directory Network (IDN.) At the time, whether to offer a controlled vocabulary search or a free-text search was resolved with a decision to support both. The feedback from the user community indicated that being asked to independently determine the appropriate 'English" words through a free-text search would be very difficult. The preference was to be 'prompted' for relevant keywords through the use of a hierarchy of well-designed science keywords. The controlled keywords serve to 'normalize' the search through knowledgeable input by metadata providers. Earth science keyword taxonomies were developed, rules for additions, deletions, and modifications were created. Secondary sets of controlled vocabularies for related descriptors such as projects, data centers, instruments, platforms, related data set link types, and locations, along with free-text searches assist users in further refining their search results. Through this robust 'search and refine' capability in the GCMD users are directed to the data and services they seek. The next step in guiding users more directly to the resources they desire is to build a 'reasoning' capability for search through the use of ontologies. Incorporating twelve sets of Earth science keyword taxonomies has boosted the GCMD S ability to help users define and more directly retrieve data of choice.

Olsen, Lola M.

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert

Increasing Data Discovery and Re-Use: The Space Life Sciences Ontology

Two of the most important goals of the adoption of the FAIR principles are increasing the ability of agents to find and re-use research data. Achieving these goals for space life sciences research is even more pressing, given the relatively expensive and scarce nature of these data. We have reported in the past on the progress made by exemplar life sciences data systems towards implementing FAIR, showing gaps particularly in the “interoperability area” of the principles; the lack of common conceptual models for space life science research is one reason for this gap. There were few available resources that define, annotate, categorize or otherwise relate various kinds of metadata describing the acquisition, nature, and intent of investigational space life sciences data. To address this gap, NASA is working with the Open Biological and Biomedical Ontology Foundry (https://obofoundry.org/) to develop the Space Life Science Ontology (SLSO) that is intended to support archival and other kinds of systems that operate using these data. The scope of the ontology includes concepts regarding those aspects of investigation design and execution specific or unique to space environments, such as types of specialized equipment, operating organizations, and documentation. The ontology is continually being developed and published to the life science community (https://github.com/nasa/LSDAO/); at the time of this publication, the SLSO newly and uniquely defines 30 types (classes), 90 properties, and 14 relations specific to space life sciences metadata. In addition, the SLSO reuses (imports) some 2,360 types (classes), 49 properties, and 393 relations from other ontologies that are relevant to these kinds of metadata. In addition to its role as a common conceptualization for space biomedical research activities, the SLSO can also be used to provide automated support for traditionally difficult and expensive activities such as data curation and cross-system data integration and analysis.

fair

Improving “Domain-Relevant Metadata Requirements” for Supporting Open-Source Science Initiative

Implementation of the NASA Open-Source Science Initiative (OSSI) requires sharing of all relevant information to ensure “open reproducible science” [1]. However, there are several challenges in applying the OSSI to airborne field campaigns focused on atmospheric composition, which often involve a wide variety of in-situ measurements for trace gases, aerosol and cloud properties, meteorological parameters, and radiation fields. To ensure open reproducibility from airborne field campaigns, it is essential to obtain detailed measurement descriptions, which include the detection principle, sample procedure and treatment, and data processing and correction method. The challenge is that some information, e.g., sampling procedure and treatment, may be instrument-specific and campaign or platform-dependent. The data processing may also involve empirical corrections which may evolve over time. In addition, these details (especially operation- or campaign-specific ones) are often not given in journal publications. Given these issues, there is a need to leverage and improve the current “domain-relevant metadata requirements” to represent the measurement description in standardized metadata. These requirements can then facilitate systematic collection of measurement specific metadata and serve as a foundation to develop tools for making the information accessible and data more interoperable and usable or reusable. Here we show a review of existing metadata collections, use cases, and needs for new standards.

Sean Leavor

Developing Data Citations from Digital Object Identifier Metadata

NASA's Earth Science Data and Information System (ESDIS) Project has been processing information for the registration of Digital Object Identifiers (DOI) for the last five years of which an automated system has been in operation for the last two years. The ESDIS DOI registration system has registered over 2000 DOIs with over 1000 DOIs held in reserve until all required information has been collected. By working towards the goal of assigning DOIs to the 8000+ data collections under its management, ESDIS has taken the first step towards facilitating the use of data citations with those products. Jeanne Behnke, ESDIS Deputy Project Manager has reviewed and approved the poster.

digital object idenifiers

The Role of Metadata Standards in EOSDIS Search and Retrieval Applications

Metadata standards play a critical role in data search and retrieval systems. Metadata tie software to data so the data can be processed, stored, searched, retrieved and distributed. Without metadata these actions are not possible. The process of populating metadata to describe science data is an important service to the end user community so that a user who is unfamiliar with the data, can easily find and learn about a particular dataset before an order decision is made. Once a good set of standards are in place, the accuracy with which data search can be performed depends on the degree to which metadata standards are adhered during product definition. NASA's Earth Observing System Data and Information System (EOSDIS) provides examples of how metadata standards are used in data search and retrieval.

Pfister, Robin

JGI Archive and Metadata Organizer (JAMO) v2.0.0

JAMO (JGI Archive and Metadata Organizer) helps researchers keep large collections of scientific files organized, findable, and safe. It lets you submit files with consistent, template-driven metadata, bundle related files into sets, and track them as a group instead of one by one. As data ages, JAMO automatically moves it from fast disk to cost-saving tape and can bring it back when needed, keeping storage lean without losing access. A simple web/CLI workflow supports submitting, checking status, retrying, and updating metadata. Compared with generic storage, JAMO's strengths are: clear, searchable metadata tuned for science; set-level organization that mirrors real projects; and built-in lifecycle care (archive, purge, restore) so you don't have to manage those steps yourself.

Cassol, Daniela [Lawrence Berkeley National Labora

Opening doors to physical sample tracking and attribution in Earth and environmental sciences

Physical samples and their associated data and metadata underpin scientific discoveries across disciplines and can enable new science when appropriately archived. However, there are significant gaps in current practices and infrastructure that prevent accurate provenance tracking, reproducibility, and attribution. For most samples, descriptive metadata are often sparse, inaccessible, or absent. Samples and associated data and metadata may also be scattered across numerous physical collections, data repositories, laboratories, data files, and papers with no clear linkage or provenance tracking as new information is generated over time. The Earth Science Information Partners (ESIP) Physical Samples Curation Cluster has therefore developed guidance for scientific authors on ‘Publishing Open Research Using Physical Samples.’ This involved synthesizing existing practices, gathering community feedback, and assessing real-world examples. We identified improvements needed to enable authors to efficiently cite and link Earth science samples and related data, and track their use. Our goal is to help improve discoverability, interoperability, and reuse of physical samples, and associated data and metadata. Though primarily focused on the needs of Earth and environmental sciences, these guidelines are broadly applicable.

58 GEOSCIENCES

ISAIA: Interoperable Systems for Archival Information Access

The ISAIA project was originally proposed in 1999 as a successor to the informal AstroBrowse project. AstroBrowse, which provided a data location service for astronomical archives and catalogs, was a first step toward data system integration and interoperability. The goals of ISAIA were ambitious: '...To develop an interdisciplinary data location and integration service for space science. Building upon existing data services and communications protocols, this service will allow users to transparently query hundreds or thousands of WWW-based resources (catalogs, data, computational resources, bibliographic references, etc.) from a single interface. The service will collect responses from various resources and integrate them in a seamless fashion for display and manipulation by the user.' Funding was approved only for a one-year pilot study, a decision that in retrospect was wise given the rapid changes in information technology in the past few years and the emergence of the Virtual Observatory initiatives in the US and worldwide. Indeed, the ISAIA pilot study was influential in shaping the science goals, system design, metadata standards, and technology choices for the virtual observatory. The ISAIA pilot project also helped to cement working relationships among the NASA data centers, US ground-based observatories, and international data centers. The ISAIA project was formed as a collaborative effort between thirteen institutions that provided data to astronomers, space physicists, and planetary scientists. Among the fruits we ultimately hoped would come from this project would be a central site on the Web that any space scientist could use to efficiently locate existing data relevant to a particular scientific question. Furthermore, we hoped that the needed technology would be general enough to allow smaller, more-focused community within space science could use the same technologies and standards to provide more specialized services. A major challenge to searching for data across a broad community is that information that describe some data products are either not relevant to other data or not applicable in the same way. Some previous metadata standard development efforts (e.g., in the earth science and library communities) have produced standards that are very large and difficult to support. To address this problem, we studied how a standard may be divided into separable pieces. Data providers that wish to participate in interoperable searches can support only those parts of the standard that are relevant to them. We prototyped a top-level metadata standard that was small and applicable to all space science data.

Hanisch, Robert J.

Evolution of Web Services in EOSDIS: Search and Order Metadata Registry (ECHO)

During 2005 through 2008, NASA defined and implemented a major evolutionary change in it Earth Observing system Data and Information System (EOSDIS) to modernize its capabilities. This implementation was based on a vision for 2015 developed during 2005. The EOSDIS 2015 Vision emphasizes increased end-to-end data system efficiency and operability; increased data usability; improved support for end users; and decreased operations costs. One key feature of the Evolution plan was achieving higher operational maturity (ingest, reconciliation, search and order, performance, error handling) for the NASA s Earth Observing System Clearinghouse (ECHO). The ECHO system is an operational metadata registry through which the scientific community can easily discover and exchange NASA's Earth science data and services. ECHO contains metadata for 2,726 data collections comprising over 87 million individual data granules and 34 million browse images, consisting of NASA s EOSDIS Data Centers and the United States Geological Survey's Landsat Project holdings. ECHO is a middleware component based on a Service Oriented Architecture (SOA). The system is comprised of a set of infrastructure services that enable the fundamental SOA functions: publish, discover, and access Earth science resources. It also provides additional services such as user management, data access control, and order management. The ECHO system has a data registry and a services registry. The data registry enables organizations to publish EOS and other Earth-science related data holdings to a common metadata model. These holdings are described through metadata in terms of datasets (types of data) and granules (specific data items of those types). ECHO also supports browse images, which provide a visual representation of the data. The published metadata can be mapped to and from existing standards (e.g., FGDC, ISO 19115). With ECHO, users can find the metadata stored in the data registry and then access the data either directly online or through a brokered order to the data archive organization. ECHO stores metadata from a variety of science disciplines and domains, including Climate Variability and Change, Carbon Cycle and Ecosystems, Earth Surface and Interior, Atmospheric Composition, Weather, and Water and Energy Cycle. ECHO also has a services registry for community-developed search services and data services. ECHO provides a platform for the publication, discovery, understanding and access to NASA s Earth Observation resources (data, service and clients). In their native state, these data, service and client resources are not necessarily targeted for use beyond their original mission. However, with the proper interoperability mechanisms, users of these resources can expand their value, by accessing, combining and applying them in unforeseen ways.

Mitchell, Andrew

Transformation of the NASA Life Sciences Portal to a FAIR Data Point

The FAIR principles emphasize optimizing metadata, the vast majority of which are textual in nature, and often organized into attribute name-value pairs. This uniformity has led to the development of guidelines and best practices for providing programmatic access to scientific data through their metadata, yielding the first iteration of the FAIR Data Point Specifications (FDPS). A key feature of the FDPS is its support for automated agents seeking and fetching data without first needing to learn a plethora of different application programming interfaces. These software agents can interrogate metadata catalogs that adhere to FDPS in a uniform manner because each catalog describes itself and its metadata schema consistently. This approach enhances the sustainability of data retrieval support, allowing systems to refine and update their metadata schemas as needed and without requiring data-seeking software agents to change how they interrogate FDPS catalogs. An essential aspect of the FDPS is the standardization of data catalog semantics, which formalizes concepts such as “metadata” and “metadata service” and links them to other concepts specifications including the Data Catalog Vocabulary (DCAT), a W3C standard that is also the basis of NASA-STD-2831 “Metadata Standard for Data Discoverability,” authored by NASA’s Office of the Chief Information Officer. The FDPS references DCAT (version 2) elements which focus on the distribution of datasets and support the goal of stream-lined catalog integration across repositories for improved data discovery. Additionally, the FDPS also prescribe the use of Linked Data Platform elements for data catalog-metadata record containment descriptions, allowing users to ascertain which data and metadata belong to which catalogs. NASA’s Life Sciences Portal is implementing the FDPS while formalizing its metadata schema to support the accelerated synthesis of knowledge from space life sciences investigations.

platform

Towards FAIR Workflows for Federated Experimental Sciences

A de-centralized, peer-to-peer AI metadata framework is demonstrated which can enable end-to-end metadata & lineage tracking for distributed Machine Learning pipelines spanning edge, High Performance Computing, and cloud environments. With a specific example of end-to-end microscopy algorithm and datasets, the proposed method shows how to enable reproducibility, audit trail, provenance of metadata artifacts. The emerging needs of automation in experimental sciences, ML-centric workflows, and FAIR metadata management across federated compute environments is addressed.

machine learning

Earth Science Informatics - Overview

Over the last 10-15 years, significant advances have been made in information management, there are an increasing number of individuals entering the field of information management as it applies to Geoscience and Remote Sensing data, and the field of informatics has come to its own. Informatics is the science and technology of applying computers and computational methods to the systematic analysis, management, interchange, and representation of science data, information, and knowledge. Informatics also includes the use of computers and computational methods to support decision making and applications. Earth Science Informatics (ESI, a.k.a. geoinformatics) is the application of informatics in the Earth science domain. ESI is a rapidly developing discipline integrating computer science, information science, and Earth science. Major national and international research and infrastructure projects in ESI have been carried out or are on-going. Notable among these are: the Global Earth Observation System of Systems (GEOSS), the European Commissions INSPIRE, the U.S. NSDI and Geospatial One-Stop, the NASA EOSDIS, and the NSF DataONE, EarthCube and Cyberinfrastructure for Geoinformatics. More than 18 departments and agencies in the U.S. federal government have been active in Earth science informatics. All major space agencies in the world, have been involved in ESI research and application activities. In the United States, the Federation of Earth Science Information Partners (ESIP), whose membership includes nearly 150 organizations (government, academic and commercial) dedicated to managing, delivering and applying Earth science data, has been working on many ESI topics since 1998. The Committee on Earth Observation Satellites (CEOS)s Working Group on Information Systems and Services (WGISS) has been actively coordinating the ESI activities among the space agencies. Remote Sensing; Earth Science Informatics, Data Systems; Data Services; Metadata

Remote Sensing; Earth Science Informatics