Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ICARTT File Format Enhancements: Supporting FAIRness and Data Discovery of Suborbital Campaign Data

Suborbital campaigns aim to accomplish a wide variety of goals and can include a variety of platforms, instruments, and parameters measured. In 2004, the ICARTT (International Consortium for Atmospheric Research on Transport and Transformation) standards were developed to fulfill data management needs for the ICARTT campaign. The ICARTT file format is text-based and composed of a header with important data description information and the data section. Built on the NASA Ames and GTE data formats, the ICARTT format was created to facilitate data exchange and promote collaborations among the science teams for achieving the ICARTT campaign goals. Due to its success and adaptation for use in many other field campaigns, the ICARTT file format became a NASA standard in 2010 and was amended in January 2017. These changes provided many enhancements, including the requirement for variable standard names. Primarily designed for airborne field studies, ICARTT has been further utilized for ground-based studies. NASA has made a commitment to build an inclusive open science community over the next decade. Open-source science strives to make publicly funded scientific research transparent, inclusive, accessible, and reproducible. The ICARTT format can host metadata that is critical for proper use of the data, particularly for in-situ measurements, and can enhance data discovery and accessibility. However, the required fields are often free text, meaning that the information is human readable, but not machine interpretable. Furthermore, the amount and type of information provided can vary significantly between principal investigators and campaigns. To support FAIR principles and interoperability, enhancements to the ICARTT standards are recommended. Possible recommendations include potential use of controlled and consistent vocabulary for variable standard name and certain common metadata elements; standardizing timestamps for easier data comparisons and analysis; and providing guidance on variable measurement units and how they are reported. Enhancing ICARTT metadata can further streamline the process to make suborbital data more readily available to the data user and improve variable-level metadata. Providing more variable-level metadata can enhance data searching and discovery, supporting NASA’s Open-Source Science Initiative (OSSI).

Megan Buzanowicz

Semantic Web Data Discovery of Earth Science Data at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC)

Mirador is a web interface for searching Earth Science data archived at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). Mirador provides keyword-based search and guided navigation for providing efficient search and access to Earth Science data. Mirador employs the power of Google's universal search technology for fast metadata keyword searches, augmented by additional capabilities such as event searches (e.g., hurricanes), searches based on location gazetteer, and data services like format converters and data sub-setters. The objective of guided data navigation is to present users with multiple guided navigation in Mirador is an ontology based on the Global Change Master directory (GCMD) Directory Interchange Format (DIF). Current implementation includes the project ontology covering various instruments and model data. Additional capabilities in the pipeline include Earth Science parameter and applications ontologies.

Hegde, Mahabaleshwara

Improvements in Space Geodesy Data Discovery at the CDDIS

The Crustal Dynamics Data Information System (CDDIS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to store space geodesy and geodynamics related data products in a central data bank. to maintain information about the archival of these data, and to disseminate these data and information in a timely manner to a global scientific research community. The archive consists of GNSS, laser ranging, VLBI, and DORIS data sets and products derived from these data. The CDDIS is one of NASA's Earth Observing System Data and Information System (EOSDIS) distributed data centers; EOSDIS data centers serve a diverse user community and arc tasked to provide facilities to search and access science data and products. Several activities are currently under development at the CDDIS to aid users in data discovery, both within the current community and beyond. The CDDIS is cooperating in the development of Geodetic Seamless Archive Centers (GSAC) with colleagues at UNAVCO and SIO. TIle activity will provide web services to facilitate data discovery within and across participating archives. In addition, the CDDIS is currently implementing modifications to the metadata extracted from incoming data and product files pushed to its archive. These enhancements will permit information about COOlS archive holdings to be made available through other data portals such as Earth Observing System (EOS) Clearinghouse (ECHO) and integration into the Global Geodetic Observing System (GGOS) portal.

Noll, C.

The Challenges of Interoperable Data Discovery

The Global Change Master Directory (GCMD) assists the oceanographic community in data discovery and access through its online metadata directory. The directory also offers data holders a means to post and search their oceanographic data through the GCMD portals, i.e. online customized subset metadata directories. The Gulf of Maine Ocean Data Partnership (GoMODP) has expressed interest in using the GCMD portals to increase the visibility of their data holding throughout the Gulf of Maine region and beyond. The purpose of the Gulf of Maine Ocean Data Partnership (GoMODP) is to "promote and coordinate the sharing, linking, electronic dissemination, and use of data on the Gulf of Maine region". The participants have decided that a "coordinated effort is needed to enable users throughout the Gulf of Maine region and beyond to discover and put to use the vast and growing quantities of data in their respective databases". GoMODP members have invited the GCMD to discuss further collaborations in view of this effort. This presentation. will focus on the GCMD GoMODP Portal - demonstrating its content and use for data discovery, and will discuss the challenges of interoperable data discovery. interoperability among metadata standards and vocabularies will be discussed. A short overview of the lessons learned at the Marine Metadata Interoperability (MMI) metadata workshop held in Boulder, Colorado on August 9-11, 2005 will be given.

Meaux, Melanie F.

Enabling Cloud Services and Enhanced Data Discovery With Earthdata-Varinfo

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. Each of these collections can contain hundreds of variables that describe each measurement within the dataset, therefore an automated method for generating UMM-Var records is necessary. The Unified Metadata Model for Variables (UMM-Var) provides a framework for variable metadata records in NASA’s Common Metadata Repository (CMR). The Python tool, earthdata-varinfo, was developed to solve this problem of automating the curation of UMM-Var records. Given either a collection DMR file or a netCDF-4 file, earthdata-varinfo can scrape variable metadata and return a CMR compliant UMM-Var record. Earthdata-varinfo can generate thousands of UMM-Var records in a matter of seconds, thus enabling subsetting capabilities and enhancing data discovery.

Eni Awowale

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov

Unified User Interface to Support Effective and Intuitive Data Discovery, Dissemination, and Analysis at NASA GES DISC

Goddard Earth Sciences Data and Information Services Center (GES DISC) has been providing access to scientific data sets since 1990s. Beginning as one of the first Earth Observing System Data and Information System (EOSDIS) archive centers, GES DISC has evolved to offer a wide range of science-enabling services. With a growing understanding of needs and goals of its science users, GES DISC continues to improve and expand on its broad set of data discovery and access tools, sub-setting services, and visualization tools. Nonetheless, the multitude of the available tools, a partial overlap of functionality, and independent and uncoupled interfaces employed by these tools often leave the end users confused as of what tools or services are the most appropriate for a task at hand. As a result, some the services remain underutilized or largely unknown to the users, significantly reducing the availability of the data and leading to a great loss of scientific productivity. In order to improve the accessibility of GES DISC tools and services, we have designed and implemented UUI, the Unified User Interface. UUI seeks to provide a simple, unified, and intuitive one-stop shop experience for the key services available at GES DISC, including sub-setting (Simple Subset Wizard), granule file search (Mirador), plotting (Giovanni), and other services. In this poster, we will discuss the main lessons, obstacles, and insights encountered while designing the UUI experience. We will also present the architecture and technology behind UUI, including NodeJS, Angular, and Mongo DB, as well as speculate on the future of the tool at GES DISC as well as in a broader context of the Space Science Informatics.

web portal

Pipeline for Applications-Based Data Discovery

From disaster response and mitigation to monitoring water quality or protecting wildlife habitat, satellite Earth observation data can be applied in countless ways to meet pressing needs and benefit society. The crucial first step toward successful data application is data discovery. Potential users often know exactly what data they need--what Earth feature or phenomenon they need to observe, how frequently, and at what resolution or level of accuracy--but may still struggle to discover the existing observations that meet their needs. We have developed a pipeline to connect applications-based users to specific satellites and data collections within NASA's Earth observation program of record that are highly relevant to their data needs. This pipeline combines available information on satellite and instrument measurement characteristics with an innovative machine learning-based approach that identifies instruments that are most relevant to the feature or phenomenon of interest.

Katrina S Virts

Scalability, Interoperability, and Security at the Data Discovery Level: A System Administrator's Perspective

The Global Change Master Directory (GCMD) has been one of the best known Earth science and global change data discovery online resources throughout its extended operational history. The growing popularity of the system since its introduction on the World Wide Web in 1994 has created an environment where resolving issues of scalability, security, and interoperability have been critical to providing the best available service to the users and partners of the GCMD. Innovative approaches developed at the GCMD in these areas will be presented with a focus on how they relate to current and future GO-ESSP community needs.

Source record

EOSDIS CMR: Shifting Data Discovery & Use into a Higher Gear

Earth observation data comes in many forms, formats, and from a multitude of sources; to make the best of a very large and diverse data catalog (data from a dozen different national Distributed Active Archive Centers (DAACs) as well as international sources), NASA has created a one-stop-shop for earth data consumers to view, find, and get the data they need regardless of its original source or format, which is powered by a Common Metadata Repository (CMR). CMR is the underpinning that allows for the visualization, search, discovery, manipulation, and acquisition of a variety of datasets, and as such it is constantly evolving to do more and serve our communities better; CMR has embraced community ownership by making itself an open-source API, being compatible with Catalog Services for the Web (CSW) and OpenSearch APIs, and by encouraging the user community to make and share improvements.

CMR

New Ways of Facilitating Improved Data Discovery and Access for NASA's Suborbital Earth Science Observations

NASA conducts field research in various Earth Science disciplines utilizing airborne and other non-satellite platforms to acquire in situ and remotely sensed observations indicative of physical processes across a range of scales. Field efforts are key in the development and validation of instruments and satellite algorithm refinements. The heterogeneous data, with a range of file formats, scales, and acquisition methods, support research in several science areas. NASA’s archive process assigns data products to discipline-oriented Distributed Active Archive Centers (DAACs) for stewardship. Over time, individual DAACs have developed tools for data browsing and serving disparate user bases. As science becomes more interdisciplinary, researchers need to incorporate observations from multiple campaigns, and multiple DAACs, into their work. Motivated in part by this shifting paradigm of needs, the Catalog of Archived Suborbital Earth Science Investigations (CASEI) was created. CASEI provides a single starting point to browse, search, and discover airborne and field data. Contextual metadata are organized and inter-linked allowing intuitive, integrated exploration across all NASA DAACs. Campaign science objectives, platform and instrument configurations, geographical details, geophysical concepts, and more are tracked in CASEI’s database, facilitating multi-parameter search, browse, and discovery of relevant data products. Researchers are able to directly access associated data products, via DOI links, regardless of the DAAC where they reside. Significant events, key time periods of high science interest within the longer-duration campaign effort, are also indicated and allow for a more efficient identification of critical data subsets. This presentation describes CASEI’s development, intensive metadata curation process, and demonstrates the web interface experience. Initial content metrics and plans for continued maintenance will also be discussed.

metadata

A Framework for Assessing Earth Observation Metadata Quality: Implications for Data Discovery and Open Science

The Common Metadata Repository (CMR) contains metadata records describing NASA’s collection of over 8,000 Earth observation data products. The Analysis and Review of CMR (ARC) Team at Marshall Space Flight Center assesses the quality of these metadata records. Metadata, rather than the data itself, is indexed for search in both discipline-specific datacenters and global or aggregated catalogs (such as Earth data Search), making it essential for determining whether a data product is appropriate for a given research question or application need. Since metadata connects users to data, it should be as accurate and complete as possible in addition to meeting minimum database requirements. The ARC team has developed a metadata quality framework by which to assess quality. The framework consists of a set of quality criteria that converge around the dimensions of correctness, completeness, and consistency, with the goal of improving the discoverability, accessibility, and usability of NASA’s Earth Observation data. The application of the framework has resulted in a measurable improvement in NASA’s metadata quality. Key aspects of the framework’s success are the ability to systematically evaluate metadata and provide actionable quality improvement recommendations. Lessons learned from the project will be shared along with implementation details which may be relevant to other science disciplines. By aiming to make data more discoverable and accessible to a broad user community, the ARC metadata quality framework helps contribute to NASA’s commitment to open science.

Jeanne Le Roux

NASA GIBS and Worldview: Leveraging Visualizations to Improve Data Discovery

NASA's Global Imagery Browse Services (GIBS) leverages scientific and community best practices and standards to provide a scalable, compliant, and authoritative source for NASA Earth Observing System (EOS) Earth science data visualizations. Since 2013, its goal has been to "transform how end users interact and discover [EOS] data through visualizations." Imagery layers within GIBS allow end users to easily and quickly interact with full resolution, pre-generated visualizations of scientific parameters. This interactive discovery approach relies on visual observation and identification of phenomena that are not as simply identified otherwise.

Global Imagery Browse Services

The Radiation Biology Ontology: A New Tool Supporting FAIR Principles Across Radiation Biology Facilitating Data Discovery and Integration

Development of the Radiation Biology Ontology (RBO) was motivated by the need for a comprehensive, well-structured ontology for encoding radiation biology metadata. The primary use-cases were archiving data in the STORE database (https://www.storedb.org/), the repository for the RadoNorm Project, and in GeneLab (https://genelab.nasa.gov), NASA’s ‘omics database. The scope of radiobiology research ranges from physics to radiation oncology to socio-legal studies; no existing ontology has the necessary breadth or depth. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR radiation biology data.

ontology

Increasing Data Discovery and Re-Use: The Space Life Sciences Ontology

Two of the most important goals of the adoption of the FAIR principles are increasing the ability of agents to find and re-use research data. Achieving these goals for space life sciences research is even more pressing, given the relatively expensive and scarce nature of these data. We have reported in the past on the progress made by exemplar life sciences data systems towards implementing FAIR, showing gaps particularly in the “interoperability area” of the principles; the lack of common conceptual models for space life science research is one reason for this gap. There were few available resources that define, annotate, categorize or otherwise relate various kinds of metadata describing the acquisition, nature, and intent of investigational space life sciences data. To address this gap, NASA is working with the Open Biological and Biomedical Ontology Foundry (https://obofoundry.org/) to develop the Space Life Science Ontology (SLSO) that is intended to support archival and other kinds of systems that operate using these data. The scope of the ontology includes concepts regarding those aspects of investigation design and execution specific or unique to space environments, such as types of specialized equipment, operating organizations, and documentation. The ontology is continually being developed and published to the life science community (https://github.com/nasa/LSDAO/); at the time of this publication, the SLSO newly and uniquely defines 30 types (classes), 90 properties, and 14 relations specific to space life sciences metadata. In addition, the SLSO reuses (imports) some 2,360 types (classes), 49 properties, and 393 relations from other ontologies that are relevant to these kinds of metadata. In addition to its role as a common conceptualization for space biomedical research activities, the SLSO can also be used to provide automated support for traditionally difficult and expensive activities such as data curation and cross-system data integration and analysis.

fair

NASA's GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate 'open science' biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics ('omics') data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

exobiology

NASAs GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate open science biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics (omics) data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

genome