Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “discoverability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

84 records · Page 5

WIS and WIGOS Metadata as the Foundation for a Sustainable Framework for Global Greenhouse Gas Watch Data Exchange

Metadata (data about data) is a critical component of data discovery, description, evaluation, documentation, and preservation. Developing and propagating metadata standards has been a longstanding area of activity in WMO and beyond. The WIS2 and WIGOS metadata models are being actively developed and maintained by dedicated task teams, established under the WMO Expert Team on Metadata. The metadata representations and vocabularies are governed by well-established processes within WMO. These standards are being used in a number of metadata/data exchange activities (e.g., WMO Information System 2.0 (WIS2), WIGOS (WMDR), Climate Data Management Systems (CMDS), etc.). It should also be noted that the application of the WIS2 and WIGOS standards fully support the WMO Unified Data Policy and open data policy as well as greatly enhance the value of observations by fostering data F.A.I.R.ness. Furthermore, the WMO metadata standards can serve as the foundation for a framework that will facilitate metadata mapping between the existing schemas used in well-established data centres, e.g., WMO WDCGG (World Data Centre for Greenhouse Gases) and NOAA ObsPack (Observation Package Data Products) and to automate metadata exchange between data centres as well as with WMO. These activities will play a central role in integrating measurements sponsored by various member countries and organizations to provide a more comprehensive characterization of the temporal and spatial distribution of the greenhouse gases. At the same time, this metadata exchange can lead to member countries and partner organizations improving their current metadata collection process for data discoverability, interoperability, and (re)usability. This presentation will describe metadata activities in the context of WIS2 and WIGOS and how they apply to GGGW data integration via metadata mapping and exchange.

Gao Chen↗

pyQuARC: Preparing for Full Release

Metadata holds the contextual information about data and is the underlying structure for many data search portals. High quality metadata optimizes search results, allowing users to quickly retrieve the data they need. With the abundant volume and diversity of Earth observation datasets, data discovery and metadata quality are critical for end users. The Common Metadata Repository (CMR), for example, currently hosts metadata for over 9,000 Earth observation data products archived across 12 NASA Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, assesses the completeness, correctness, and consistency of these metadata records to ensure they are accessible, usable, and discoverable. In 2021, ARC began developing pyQuARC, an open source library for Earth Observation Metadata Quality Assessment to automate this effort. The tool uses ARC’s existing metadata quality framework to provide prioritized recommendations for metadata improvement. During initial testing, pyQuARC automatically identified 58% of metadata findings when compared with a sample of manually reviewed records. Using the results from initial testing, this presentation will focus on recent advancements and improvements of the tool as the ARC team prepares for pyQuARC’s full release. It will also demonstrate pyQuARC's enrichment value, not only for the ARC team, but the broader EOSDIS metadata community as well.

Essence Raphael↗

The TOLNet 2.0 Website: How an API Can Promote Open Science and FAIR Principles

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and other ozone relevant research. To support the broad spectrum of data use, the TOLNet team launched a major effort to upgrade the web-based data repository aiming to enhance the data discoverability and to enable machine-to-machine data upload and download processes. Specifically, the TOLNet website included an application programming interface (API), which supports machine-to-machine data search and data download. The API also extracts selected variables from the files, which can be retrieved as JSON objects and used to create data displays without having to download or open the underlying files. The TOLNet science team members can also use the API for automated data upload, including a data file scanning feature to ensure data product integrity. To be presented will include a summary of key features of data repositories, an actual use case of machine-to-machine data access/use, as well as our journey to make TOLNet data more FAIR, i.e., more findable, accessible, interoperable, and (re)usable.

Crystal Gummo↗

Breaking Barriers: Integrating Geo-Leo Aerosol Data with an Open-Source Approach

The scientific community is still examining the novel data from geostationary satellite observations and evaluating methods for effectively fusing the polar observations with various spatial and temporal resolutions. However, the merged data will present a significant ""Big Data"" challenge, including processing, storage, data discoverability, accessibility, and migration within cloud computing environments. We have developed an open-source package to fuse aerosol optical depths (AOD) products from six satellite sensors in the past four years (2019~2023), and this presentation will update our recent progress. Using this Python-based package, we produced a level 3 global (AOD) product in a quarter-degree spatial resolution every half-hour, fusing the Level 2 AOD data with the Dark Target aerosol retrieval algorithm from six satellites: three geostationary (GOES-16/17 and Himawari-8) with high temporal resolution, and three polar orbiting (TERRA/MODIS, AQUA/MODIS, and SNPP-VIIRS) with global coverage. By integrating these observations, the diurnal cycle of global AOD in this fused product can be characterized at local, regional, and global scales. Furthermore, we are committed to openness and transparency by providing our package and its associated functionalities as open-source. Our dedication to adhering to the FAIR, CARE, and TRUST principles ensures that our users can rely on the integrity and ethical standards of our work. For instance of Interoperability, this package fuses remote sensing products on demand into desired temporal and spatial domains. It can be run in a central processing unit (CPU) or a Graphics processing unit (GPU) mode. This package will empower researchers and practitioners to use satellite and sensor data efficiently in various applications and research.

Xiaohua Pan↗

Informing Wildfire Needs: The Expanded User Interface of NASA's Fire Information for Resource Management System (Firms)

As the global community continues to experience, and respond to, living in a changing environment, access to tools, technologies, and timely data utilized by an increasingly diverse set of stakeholders is increasing. This year, 2023, has thus far seen an unprecedented number of extreme events, increasingly driven by changes in the climate and a strong 2023 ENSO pattern. In Canada, a record number of wildfires, and associated weather events, evacuations, and infrastructure and habitat destruction has taken place, and is ongoing. Large swaths of Greece have experienced similar wildfire destruction. Most recently, Maui has experienced destructive wildfires, and early 2023 saw massive wildfires in Chile. NASA's Fire Information for Resource Management System, or FIRMS, has a fifteen-year history of providing timely and comprehensive data and information on wildfires to stakeholders. FIRMS was initially developed in 2007 by the University of Maryland, with funds from NASA's Applied Sciences Program and the United Nations Food and Agriculture Organization (UN FAO), to provide near real-time active fire Locations to natural resource managers that faced challenges obtaining timely satellite-derived fire information. FIRMS has consistently evolved to address the needs of stakeholders Living in a changing environment; in 2012 it transitioned to NASA LANCE and in 2021 through a partnership between NASA and the US Forest Service, an updated version of FIRMS was released for the US and Canada. As NASA and other federal agencies continue to accelerate Open Science through integrated efforts such as the Year of Open Science, the provision of readily discoverable, findable, accessible, interoperable, reusable data represents a major focus to facilitate equitable outcomes. FIRMS supports this acceleration in Open Science by continuing to provision data and information for its traditional user base, while addressing the novel user needs of an increasingly diverse set of stakeholders seeking robust, reliable, transparently generated data and information. Increasingly, FIRMS is utilized by citizen scientists and individuals directly affected by wildfires - through evacuations, risks to structures/homes, poor air quality. etc. FIRMS has also been Leveraged to detect and assess the impacts resulting from ongoing conflicts. This further highlights the multi-faceted impacts of wildfires and other events. In the Fall of 2023, FIRMS will release an expanded User Interface (UI). This interface captures and reflects the needs of, and input from, a multitude of users. These users range from federal agency representatives to non-government organizations to the private sector to citizen science entities. To respond to this expansive and diverse user need base, the updated FIRMS UI will capture a range of features to support those beginning to explore the range of data and tools available to inform wildfire awareness and knowledge. These users are supported through a Basic Mode interface, furnishing access to a light set of functionalities that provision straight-forward, readily usable information and data, and ingestible knowledge. The Advanced Mode interface supports those stakeholder groups already proficient in navigating FIRMS. These stakeholders, representing fire managers and others, perform active fire management and tactical wildfire response activities. For these stakeholders, additional datasets have been included which require in-depth knowledge of both the utility as well as the caveats of such datasets. Additional functionalities have also been embedded to aid specific user queries. The expanded UI will introduce a new Experimental Mode. The focus of this UI will be to support the provision of emerging and innovative datasets that are in development for review and comment by the user community Examples include post-fire products generated by NASA's Earth Information System (EIS) Fire. This presentation will provide an overview of the expanded FIRMS UI. We will discuss how this UI is designed to be scalable and support the unique needs of an expanding and diverse user base. We will highlight key features, elements, and datasets, and describe how user needs have informed and guided the design of the UI. We will also share recent use cases to convey, and increase awareness, among conference participants. As the global community faces more extreme wildfires, due to climate variability and change, there is an increased need for reliable data to inform, manage, and mitigate the impacts of these events. Through this work, NASA FIRMS is striving to level the playing field, by making information accessible to all; from policy makers to the private sector to historically marginalized communities. In doing so, NASA is promoting the all-hands-on-deck response needed to minimize the impacts of wildfires and harness the strengths of open science to address the greatest environmental challenge faced.

Jenny Hewson↗

The Science Discovery Engine: Connecting Heterogeneous Scientific Data and Information

Transformative science often occurs at the boundaries of different disciplines. Making interdisciplinary science data, software and documentation discoverable and accessible is essential to enabling transformative science. However, connecting this diverse and heterogeneous information is often a challenge due to several factors including the dispersed and sometimes isolated nature of data and the semantic differences between topical areas. NASA’s Science Discovery Engine (SDE) has developed several approaches to tackling these challenges. The SDE is a unified, insightful search experience that enables discovery of NASA’s open science data across five topical areas: astrophysics, biological and physical sciences, Earth science, heliophysics and planetary science. In this presentation, we will discuss our efforts to develop a systematic scientific curation workflow to integrate diverse content into a single search environment. We will also share lessons learned from our work to create a metadata crosswalk across the five disciplines.

Kaylin Bugbee↗

Transformation of the NASA Life Sciences Portal to a FAIR Data Point

The FAIR principles emphasize optimizing metadata, the vast majority of which are textual in nature, and often organized into attribute name-value pairs. This uniformity has led to the development of guidelines and best practices for providing programmatic access to scientific data through their metadata, yielding the first iteration of the FAIR Data Point Specifications (FDPS). A key feature of the FDPS is its support for automated agents seeking and fetching data without first needing to learn a plethora of different application programming interfaces. These software agents can interrogate metadata catalogs that adhere to FDPS in a uniform manner because each catalog describes itself and its metadata schema consistently. This approach enhances the sustainability of data retrieval support, allowing systems to refine and update their metadata schemas as needed and without requiring data-seeking software agents to change how they interrogate FDPS catalogs. An essential aspect of the FDPS is the standardization of data catalog semantics, which formalizes concepts such as “metadata” and “metadata service” and links them to other concepts specifications including the Data Catalog Vocabulary (DCAT), a W3C standard that is also the basis of NASA-STD-2831 “Metadata Standard for Data Discoverability,” authored by NASA’s Office of the Chief Information Officer. The FDPS references DCAT (version 2) elements which focus on the distribution of datasets and support the goal of stream-lined catalog integration across repositories for improved data discovery. Additionally, the FDPS also prescribe the use of Linked Data Platform elements for data catalog-metadata record containment descriptions, allowing users to ascertain which data and metadata belong to which catalogs. NASA’s Life Sciences Portal is implementing the FDPS while formalizing its metadata schema to support the accelerated synthesis of knowledge from space life sciences investigations.

platform↗

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

CALIPSO Final Data Product Status

The CALIPSO project is preparing for the end of the mission in October 2025. In this poster we provide detailed summaries of each final data products that will be released, a concise description of the changes between each of these final products and their previous versions, and a target release schedule. A summary of all of the steps that will be carried outby the CALIPSO project during the remainder of the mission for long term data discoverability and accessibility will also be provided.

Brian Getzewich↗

GES DISC Data Recipes in Jupyter Notebooks

The Earth Science Data and Information System (ESDIS) Project manages twelve Distributed Active Archive Centers (DAACs) which are geographically dispersed across the United States. The DAACs are responsible for ingesting, processing, archiving, and distributing Earth science data produced from various sources (satellites, aircraft, field measurements, etc.). In response to projections of an exponential increase in data production, there has been a recent effort to prototype various DAAC activities in the cloud computing environment. This, in turn, led to the creation of an initiative, called the Cloud Analysis Toolkit to Enable Earth Science (CATEES), to develop a Python software package in order to transition Earth science data processing to the cloud. This project, in particular, supports CATEES and has two primary goals. One, to transition data recipes created by the Goddard Earth Science Data and Information Service Center (GES DISC) into an interactive and educational environment using JupyterNotebooks. Two, to acclimate Earth scientists to cloud computing. To accomplish these goals, we create JupyterNotebooks to compartmentalize the different steps of data analysis and help users obtain and parse data from the command line. We also develop a Docker container, comprised of Jupyter Notebooks, Python dependencies, and command line tools, and configure it into an easy-to-deploy package. The end result is an end-to-end product that simulates the use case of end users working in the cloud computing environment.

discoverability↗