Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Lowering Barriers to Science and Space Weather Research at the Community Coordinated Modeling Center (CCMC)

The Space Weather and Heliophysics research and modeling community has been pushing the limits of our ability to understand and predict space weather events. The Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) supports the community by providing a convenient collaborative platform hosting space weather models, model simulation data, curated datasets of solar events, and associated value-added services. Using these services, researchers and other end-users may exercise, evaluate, and intercompare contributed models, triage designated R2O models, as well as collaborate on a continuously updated archive of model run results. We will focus on CCMC’s ongoing commitment to the principles and guidelines of the Open Science initiative. Particularly, we will discuss our work towards making our services more transparent and our library of model simulations more accessible, open, and reproducible. We will introduce our recent tools for data discovery and correlative analysis designed to further increase the value of the user-generated data and metadata. We will also present our recent work on making heliophysical models more accessible and open to the community, particularly through simplified user experience and expert domain support. We will report on our progress in establishing an inter-center infrastructure with the ESA Virtual Space Weather Modelling Centre (VSWMC), designed to cross organizational boundaries and provide streamlined access to a joint palette of the models.

space weather↗

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database↗

NASA Global Satellite and Model Data Products and Services for Tropical Meteorology and Climatology

Satellite remote sensing and model data play an important role in research and applications of tropical meteorology and climatology over vast, data sparse oceans and remote continents. Since the first weather satellite was launched by NASA in 1960, a large collection of NASA's Earth science data is freely available to the research and application communities around the world, significantly improving our overall understanding of the Earth system and environment. Established in the mid-80s, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), located in Maryland, USA, is a data archive center for multidisciplinary, satellite and model assimilation data products. As one of the 12 NASA data centers in Earth Sciences, GES DISC hosts several important NASA satellite missions for tropical meteorology and climatology such as the Tropical Rainfall Measuring Mission (TRMM), the Global Precipitation Measurement (GPM) mission and the Modern-Era Retrospective analysis for Research and Applications (MERRA). Over the years, GES DISC has developed data services to facilitate data discovery, access, distribution, analysis and visualization, including Giovanni, an online analysis and visualization tool without the need to download data and software. Despite many efforts for improving data access, still quite a number of challenges remain, such as finding datasets and services for a specific research topic or project, especially for inexperienced users or users outside the remote sensing community. In this article, we list and describe major NASA satellite remote sensing and model datasets and services for tropical meteorology and climatology along with examples of using the data and services, in hope that may help users better utilize the information in their research and applications.

tropical meteorology and climatology, data, servic↗

A model for live mission data systems using the OAIS reference model

Space sciences are confronted with overwhelming volume of data. The data rates are increasing, the granularity of registered observations is continuously refining, and computer technology allows producing terabytes of images and catalogs. The inexpensive emerging storage technologies, combined with the availability of high-speed communications will offer the infrastructure for extremely large data repositories to be accessible on-line. Mission data will be quickly accessible almost immediately after it has been collected from space observations. On-line science will demand for new tools and technologies for data access, data analysis, and data discovery. These trends will enhance the archival operational concepts mainly related to the long-term information preservation, placing an equally important emphasis on rapid data production, and dissemination to consumers.

mission data systems srchive system OAIS data mana↗

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian↗

pyQuARC: Preparing for Full Release

Metadata holds the contextual information about data and is the underlying structure for many data search portals. High quality metadata optimizes search results, allowing users to quickly retrieve the data they need. With the abundant volume and diversity of Earth observation datasets, data discovery and metadata quality are critical for end users. The Common Metadata Repository (CMR), for example, currently hosts metadata for over 9,000 Earth observation data products archived across 12 NASA Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, assesses the completeness, correctness, and consistency of these metadata records to ensure they are accessible, usable, and discoverable. In 2021, ARC began developing pyQuARC, an open source library for Earth Observation Metadata Quality Assessment to automate this effort. The tool uses ARC’s existing metadata quality framework to provide prioritized recommendations for metadata improvement. During initial testing, pyQuARC automatically identified 58% of metadata findings when compared with a sample of manually reviewed records. Using the results from initial testing, this presentation will focus on recent advancements and improvements of the tool as the ARC team prepares for pyQuARC’s full release. It will also demonstrate pyQuARC's enrichment value, not only for the ARC team, but the broader EOSDIS metadata community as well.

Essence Raphael↗

ESIP Documentation Cluster Session: GCMD Keyword Update

The Global Change Master Directory (GCMD) Keywords are a hierarchical set of controlled Earth Science vocabularies that help ensure Earth science data and services are described in a consistent and comprehensive manner and allow for the precise searching of collection-level metadata and subsequent retrieval of data and services. Initiated over twenty years ago, the GCMD Keywords are periodically analyzed for relevancy and will continue to be refined and expanded in response to user needs. This talk explores the current status of the GCMD keywords, the value and usage that the keywords bring to different tools/agencies as it relates to data discovery, and how the keywords relate to SWEET (Semantic Web for Earth and Environmental Terminology) Ontologies.

data discover↗

NASA GIBS and Worldview: Bringing 20 Years of Terra Data into View

For nearly 10 years, the NASA Global Imagery Browse Services (GIBS) and Worldview interactive mapping site have provided users full resolution visualizations of Terra land, cryosphere, ocean, and atmosphere science parameters available within hours of acquisition. Over that time, GIBS and Worldview have expanded their Terra visualization suite to include even more parameters covering the entire mission of science data. Users can now view daily visualizations of nearly 50 near-real time and over 125 science quality science parameters from Terra instruments. These visualizations provide a unique capability for a broad user base to interact and discover the wealth of information sense by instruments on the Terra platform. By viewing Terra-based visualizations in a single location, users can correlate retrievals across instruments. Additionally, GIBS and Worldview include visualizations from many other platforms, allowing for an even greater data discovery capability.The GIBS and Worldview teams continue to work with Terra instrument science teams to add more visualized layers, as well as supporting visualization updates as the data is continually improved. Additionally, the GIBS interfaces and Worldview site are actively working on improved visualization functionality, including vector- and granule/swath-based products, to better address the expanding needs of its user community. This presentation will focus on an overview of existing and future visualization products and capabilities that have supported, and will support, a broad use of Terra data in science, media, and application communities.

Cechini, Matthew↗

Exploiting Dark Information Resources to Create New Value Added Services to Study Earth Science Phenomena

This paper presents two research applications exploiting unused metadata resources in novel ways to aid data discovery and exploration capabilities. The results based on the experiments are encouraging and each application has the potential to serve as a useful standalone component or service in a data system. There were also some interesting lessons learned while designing the two applications and these are presented next.

Earth Science Informatics↗

NASA Giovanni: Analyze, Compare, and Visualize 2000+ Earth Satellite and Model Variables Without Downloading Data and Software

Over vast oceans and remote continents, observations are often scarce and discontinuous. Satellite and model data play a critical role in research and applications. However, finding and accessing satellite and model data can be a daunting task for many, especially those outside the community. The NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), one of 12 NASA Science Mission Directorate Data Centers, provides Earth science data, information, and services to everyone such as researchers, application users, educators, and students. GES DISC archives and supports datasets applicable to several NASA Earth Science Focus Areas including Atmospheric Composition, Water & Energy Cycles, Carbon Cycle & Ecosystem, and Climate Variability. To facilitate data discovery, evaluation, and exploration, GES DISC has developed the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni), an online tool to analyze and visualize NASA remote sensing and model data without downloading data and software. As of this writing, over 2000 Earth satellite and model variables are available in Giovanni, including several wellknown NASA satellite missions (e.g., TRMM, GPM) and projects (e.g., MERRA-2, GPCP). Giovanni provides twenty-two plots that can be used to analyze, compare, and explore Earth data across different disciplines. Results can be shared with colleagues and downloaded for further analysis. Over the years, Giovanni has helped publish over 3000 referral papers. In this presentation, we will showcase key variables and plot types in Giovanni with examples. In particular, we will present several popular precipitation products from GPM and CPCP for evaluation and comparison.

data analysis↗

Managing and Servicing Physical Oceanographic Data at a NASA Distributed Active Archive Center

The NASA Earth Science Data Information Systems Project funds and operates 12 Distributed Active Archive Center(s) (DAAC) throughout the United States. Of these 12 centers, the Physical Oceanography DAAC (PO.DAAC) is committed to providing long term archival, distribution and stewardship for NASA physical oceanographic data, primarily derived from space-born satellite systems, but also including a growing set of recent and future in situ observations from the SPURS-1 and SPURS-2 campaigns. Notable NASA missions supported include: Seasat, TOPEX/Poseidon, NSCAT, QuikSCAT, ISS-RapidScat, Jason-1, Jason-2/OSTM, GRACE, Aquarius, GHRSST, and MODIS. The following interagency and international missions are also supported by PO.DAAC: AVHRR, Coriolis, DMSP, MetOp-A, MetOp-B, Oceansat-2. The PO.DAAC currently holds 525 datasets in public distribution, spanning the following observational parameters: sea surface temperature, sea surface salinity, ocean color, ocean surface currents, ocean surface wind speed, ocean surface wind direction, sea surface height, significant wave height, ocean water mass/thickness, and sea ice age. A hundred of these datasets are available in near-real-time. Datasets are distributed through a variety of open-source access protocols including FTP, OPeNDAP, and THREDDS. FTP will soon be phased out in favor of a recently introduced HTTPS PO.DAAC Drive interface that supports WebDAV and interoperable machine-to-machine communication. OPeNDAP supports remote data/metadata query, subset, and download. THREDDS provides the features of OPeNDAP with the additional feature of temporal aggregation. PO.DAAC also offers proprietary tools and services to further enhance the data discovery, visualization and analysis experience, including but not limited to: State of the Ocean, Web Services (data/metadata discovery and extraction), HiTIDE Level-2 subsetter, Live Access Server (LAS), Webification (w10nsci), and Rich Site Summary (RSS) Datacasting. To assist with provenance of datasets, PO.DAAC has implemented DOIs for the data it distributes so that they can be properly cited. There is a user forum and helpdesk that contains data recipes and via which users can get guidance. In summary, this presentation aims to provide a general overview of PO.DAAC’s web portal and data holdings along with a set of illustrative examples leading prospective data users into the practical utility of its tools and services.

Moroni, David F.↗

NASA R&M Efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) Digital Assets

At the intersection of mission, technology, and place is NASA’s need to modernize for a digital-forward future. Digitalization, the process of moving toward digital business, is occurring everywhere and remains an ongoing process across the federal government.”[1] Whereas, Digital Transformation is “employing digitization/digital technologies (e.g., Artificial Intelligence (AI), mobile, cloud, data) to change a process, product, or capability so dramatically (e.g., real-time, intelligent, personalized, anywhere, anytime) that it is unrecognizable compared to its traditional form.” [2] In order to facilitate a digital transformation it is essential for NASA to understand and identify where data exists today and which data are value-needed in the future, understand where there are unfulfilled data needs that limit the advancement of NASA work, and ensure NASA efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) digital assets in the future. Therefore, NASA’s Reliability & Maintainability (R&M) Enterprise Data Sharing team is working to leverage both Digitization and Digital Transformation to achieve their vision of developing an R&M data discovery framework that enables our community, our partners, and our stakeholders with the ability to efficiently, robustly, and seamlessly access information that enables real-time knowledge and model-based, analytics driven, decision-making impacting R&M. As a result the R&M Enterprise Data Sharing team has conducted a survey of its Reliability, Maintainability, and Availability (RMA) community members to identify data existence (created or used) and where there are corresponding barriers to data acquisition and/or R&M or other issues as shown within this presentation.

Digital Transformation↗

Trilateral Task Force – Reliability Analysis Supporting Mission Extension/Post Mission Disposal

At the intersection of mission, technology, and place is NASA’s need to modernize for a digital-forward future. Digitalization, the process of moving toward digital business, is occurring everywhere and remains an ongoing process across the federal government.”[1] Whereas, Digital Transformation is “employing digitization/digital technologies (e.g., Artificial Intelligence (AI), mobile, cloud, data) to change a process, product, or capability so dramatically (e.g., real-time, intelligent, personalized, anywhere, anytime) that it is unrecognizable compared to its traditional form.” [2] In order to facilitate a digital transformation it is essential for NASA to understand and identify where data exists today and which data are value-needed in the future, understand where there are unfulfilled data needs that limit the advancement of NASA work, and ensure NASA efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) digital assets in the future. Therefore, NASA’s Reliability & Maintainability (R&M) Enterprise Data Sharing team is working to leverage both Digitization and Digital Transformation to achieve their vision of developing an R&M data discovery framework that enables our community, our partners, and our stakeholders with the ability to efficiently, robustly, and seamlessly access information that enables real-time knowledge and model-based, analytics driven, decision-making impacting R&M. As a result the R&M Enterprise Data Sharing team has conducted a survey of its Reliability, Maintainability, and Availability (RMA) community members to identify data existence (created or used) and where there are corresponding barriers to data acquisition and/or R&M or other issues as shown within this presentation.

Digital Transformation, Reliability Engineering↗

A Global Repository for Planet-Sized Experiments and Observations

Working across U.S. federal agencies, international agencies, and multiple worldwide data centers, and spanning seven international network organizations, the Earth System Grid Federation (ESGF) allows users to access, analyze, and visualize data using a globally federated collection of networks, computers, and software. Its architecture employs a system of geographically distributed peer nodes that are independently administered yet united by common federation protocols and application programming interfaces (APIs). The full ESGF infrastructure has now been adopted by multiple Earth science projects and allows access to petabytes of geophysical data, including the Coupled Model Intercomparison Project (CMIP) output used by the Intergovernmental Panel on Climate Change assessment reports. Data served by ESGF not only include model output (i.e., CMIP simulation runs) but also include observational data from satellites and instruments, reanalyses, and generated images. Metadata summarize basic information about the data for fast and easy data discovery.

Earth Systems Grid Federation (ESGFC)↗

The I4 Online Query Tool for Earth Observations Data

The NASA Earth Observation System Data and Information System (EOSDIS) delivers an average of 22 terabytes per day of data collected by orbital and airborne sensor systems to end users through an integrated online search environment (the Reverb/ECHO system). Earth observations data collected by sensors on the International Space Station (ISS) are not currently included in the EOSDIS system, and are only accessible through various individual online locations. This increases the effort required by end users to query multiple datasets, and limits the opportunity for data discovery and innovations in analysis. The Earth Science and Remote Sensing Unit of the Exploration Integration and Science Directorate at NASA Johnson Space Center has collaborated with the School of Earth and Space Exploration at Arizona State University (ASU) to develop the ISS Instrument Integration Implementation (I4) data query tool to provide end users a clean, simple online interface for querying both current and historical ISS Earth Observations data. The I4 interface is based on the Lunaserv and Lunaserv Global Explorer (LGE) open-source software packages developed at ASU for query of lunar datasets. In order to avoid mirroring existing databases - and the need to continually sync/update those mirrors - our design philosophy is for the I4 tool to be a pure query engine only. Once an end user identifies a specific scene or scenes of interest, I4 transparently takes the user to the appropriate online location to download the data. The tool consists of two public-facing web interfaces. The Map Tool provides a graphic geobrowser environment where the end user can navigate to an area of interest and select single or multiple datasets to query. The Map Tool displays active image footprints for the selected datasets (Figure 1). Selecting a footprint will open a pop-up window that includes a browse image and a link to available image metadata, along with a link to the online location to order or download the actual data. Search results are either delivered in the form of browse images linked to the appropriate online database, similar to the Map Tool, or they may be transferred within the I4 environment for display as footprints in the Map Tool. Datasets searchable through I4 (http://eol.jsc.nasa.gov/I4_tool) currently include: Crew Earth Observations (CEO) cataloged and uncataloged handheld astronaut photography; Sally Ride EarthKAM; Hyperspectral Imager for the Coastal Ocean (HICO); and the ISS SERVIR Environmental Research and Visualization System (ISERV). The ISS is a unique platform in that it will have multiple users over its lifetime, and that no single remote sensing system has a permanent internal or external berth. The open source I4 tool is designed to enable straightforward addition of new datasets as they become available such as ISS-RapidSCAT, Cloud Aerosol Transport System (CATS), and the High Definition Earth Viewing (HDEV) system. Data from other sensor systems, such as those operated by the ISS International Partners or under the auspices of the US National Laboratory program, can also be added to I4 provided sufficient access to enable searching of data or metadata is available. Commercial providers of remotely sensed data from the ISS may be particularly interested in I4 as an additional means of directing potential customers and clients to their products.

Stefanov, William L.↗

Integrated support of NASA satellite and in situ oceanographic data via the PO.DAAC

The NASA Physical Oceanography DAAC (PO.DAAC) serves as one of the premier repositories for oceanographic satellite data. More recently, however, it is also increasingly archiving and distributing complementary in situ datasets from NASA-sponsored field campaigns. Here we present an overview of these projects, the complex multivariate data they produce, and some of the data interoperability challenges faced when dealing with such a heterogeneous suite of observations. We summarize the range of online tools and services currently available via the PODAAC, including data discovery services, web-services for subsetting/extraction, compliance checking and visualization. We also preview some new capabilities under development that may feature in future. These efforts are indicative of an evolution of DAAC services in pursuit of our broader vision: a more integrated approach to multi-sensor oceanographic data access and delivery spanning NASA satellite missions and field campaigns in support of science and applications for societal benefit.

Vannan, Suresh↗

Collaborative Data Publication Utilizing the Open Data Repository's (ODR) Data Publisher

Introduction: For small communities in diverse fields such as astrobiology, publishing and sharing data can be a difficult challenge. While large, homogenous fields often have repositories and existing data standards, small groups of independent researchers have few options for publishing standards and data that can be utilized within their community. In conjunction with teams at NASA Ames and the University of Arizona, the Open Data Repository's (ODR) Data Publisher has been conducting ongoing pilots to assess the needs of diverse research groups and to develop software to allow them to publish and share their data collaboratively. Objectives: The ODR's Data Publisher aims to provide an easy-to-use and implement software tool that will allow researchers to create and publish database templates and related data. The end product will facilitate both human-readable interfaces (web-based with embedded images, files, and charts) and machine-readable interfaces utilizing semantic standards. Characteristics: The Data Publisher software runs on the standard LAMP (Linux, Apache, MySQL, PHP) stack to provide the widest server base available. The software is based on Symfony (www.symfony.com) which provides a robust framework for creating extensible, object-oriented software in PHP. The software interface consists of a template designer where individual or master database templates can be created. A master database template can be shared by many researchers to provide a common metadata standard that will set a compatibility standard for all derivative databases. Individual researchers can then extend their instance of the template with custom fields, file storage, or visualizations that may be unique to their studies. This allows groups to create compatible databases for data discovery and sharing purposes while still providing the flexibility needed to meet the needs of scientists in rapidly evolving areas of research. Research: As part of this effort, a number of ongoing pilot and test projects are currently in progress. The Astrobiology Habitable Environments Database Working Group is developing a shared database standard using the ODR's Data Publisher and has a number of example databases where astrobiology data are shared. Soon these databases will be integrated via the template-based standard. Work with this group helps determine what data researchers in these diverse fields need to share and archive. Additionally, this pilot helps determine what standards are viable for sharing these types of data from internally developed standards to existing open standards such as the Dublin Core (http://dublincore.org) and Darwin Core (http://rs.twdg.org) metadata standards. Further studies are ongoing with the University of Arizona Department of Geosciences where a number of mineralogy databases are being constructed within the ODR Data Publisher system. Conclusions: Through the ongoing pilots and discussions with individual researchers and small research teams, a definition of the tools desired by these groups is coming into focus. As the software development moves forward, the goal is to meet the publication and collaboration needs of these scientists in an unobtrusive and functional way.

easy to use and implement software tool↗