Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Exploring and Analyzing Climate Variations Online by Using NASA MERRA-2 Data at GES DISC

NASA Giovanni (Goddard Interactive Online Visualization ANd aNalysis Infrastructure) (http:giovanni.sci.gsfc.nasa.govgiovanni) is a web-based data visualization and analysis system developed by the Goddard Earth Sciences Data and Information Services Center (GES DISC). Current data analysis functions include Lat-Lon map, time series, scatter plot, correlation map, difference, cross-section, vertical profile, and animation etc. The system enables basic statistical analysis and comparisons of multiple variables. This web-based tool facilitates data discovery, exploration and analysis of large amount of global and regional remote sensing and model data sets from a number of NASA data centers. Long term global assimilated atmospheric, land, and ocean data have been integrated into the system that enables quick exploration and analysis of climate data without downloading, preprocessing, and learning data. Example data include climate reanalysis data from NASA Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) which provides data beginning in 1980 to present; land data from NASA Global Land Data Assimilation System (GLDAS), which assimilates data from 1948 to 2012; as well as ocean biological data from NASA Ocean Biogeochemical Model (NOBM), which provides data from 1998 to 2012. This presentation, using surface air temperature, precipitation, ozone, and aerosol, etc. from MERRA-2, demonstrates climate variation analysis with Giovanni at selected regions.

knowledge base

Lowering Barriers to Science and Space Weather Research at the Community Coordinated Modeling Center (CCMC)

The Space Weather and Heliophysics research and modeling community has been pushing the limits of our ability to understand and predict space weather events. The Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) supports the community by providing a convenient collaborative platform hosting space weather models, model simulation data, curated datasets of solar events, and associated value-added services. Using these services, researchers and other end-users may exercise, evaluate, and intercompare contributed models, triage designated R2O models, as well as collaborate on a continuously updated archive of model run results. We will focus on CCMC’s ongoing commitment to the principles and guidelines of the Open Science initiative. Particularly, we will discuss our work towards making our services more transparent and our library of model simulations more accessible, open, and reproducible. We will introduce our recent tools for data discovery and correlative analysis designed to further increase the value of the user-generated data and metadata. We will also present our recent work on making heliophysical models more accessible and open to the community, particularly through simplified user experience and expert domain support. We will report on our progress in establishing an inter-center infrastructure with the ESA Virtual Space Weather Modelling Centre (VSWMC), designed to cross organizational boundaries and provide streamlined access to a joint palette of the models.

space weather

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database

NASA Global Satellite and Model Data Products and Services for Tropical Meteorology and Climatology

Satellite remote sensing and model data play an important role in research and applications of tropical meteorology and climatology over vast, data sparse oceans and remote continents. Since the first weather satellite was launched by NASA in 1960, a large collection of NASA's Earth science data is freely available to the research and application communities around the world, significantly improving our overall understanding of the Earth system and environment. Established in the mid-80s, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), located in Maryland, USA, is a data archive center for multidisciplinary, satellite and model assimilation data products. As one of the 12 NASA data centers in Earth Sciences, GES DISC hosts several important NASA satellite missions for tropical meteorology and climatology such as the Tropical Rainfall Measuring Mission (TRMM), the Global Precipitation Measurement (GPM) mission and the Modern-Era Retrospective analysis for Research and Applications (MERRA). Over the years, GES DISC has developed data services to facilitate data discovery, access, distribution, analysis and visualization, including Giovanni, an online analysis and visualization tool without the need to download data and software. Despite many efforts for improving data access, still quite a number of challenges remain, such as finding datasets and services for a specific research topic or project, especially for inexperienced users or users outside the remote sensing community. In this article, we list and describe major NASA satellite remote sensing and model datasets and services for tropical meteorology and climatology along with examples of using the data and services, in hope that may help users better utilize the information in their research and applications.

tropical meteorology and climatology, data, servic

A model for live mission data systems using the OAIS reference model

Space sciences are confronted with overwhelming volume of data. The data rates are increasing, the granularity of registered observations is continuously refining, and computer technology allows producing terabytes of images and catalogs. The inexpensive emerging storage technologies, combined with the availability of high-speed communications will offer the infrastructure for extremely large data repositories to be accessible on-line. Mission data will be quickly accessible almost immediately after it has been collected from space observations. On-line science will demand for new tools and technologies for data access, data analysis, and data discovery. These trends will enhance the archival operational concepts mainly related to the long-term information preservation, placing an equally important emphasis on rapid data production, and dissemination to consumers.

mission data systems srchive system OAIS data mana

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian

pyQuARC: Preparing for Full Release

Metadata holds the contextual information about data and is the underlying structure for many data search portals. High quality metadata optimizes search results, allowing users to quickly retrieve the data they need. With the abundant volume and diversity of Earth observation datasets, data discovery and metadata quality are critical for end users. The Common Metadata Repository (CMR), for example, currently hosts metadata for over 9,000 Earth observation data products archived across 12 NASA Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, assesses the completeness, correctness, and consistency of these metadata records to ensure they are accessible, usable, and discoverable. In 2021, ARC began developing pyQuARC, an open source library for Earth Observation Metadata Quality Assessment to automate this effort. The tool uses ARC’s existing metadata quality framework to provide prioritized recommendations for metadata improvement. During initial testing, pyQuARC automatically identified 58% of metadata findings when compared with a sample of manually reviewed records. Using the results from initial testing, this presentation will focus on recent advancements and improvements of the tool as the ARC team prepares for pyQuARC’s full release. It will also demonstrate pyQuARC's enrichment value, not only for the ARC team, but the broader EOSDIS metadata community as well.

Essence Raphael

ESIP Documentation Cluster Session: GCMD Keyword Update

The Global Change Master Directory (GCMD) Keywords are a hierarchical set of controlled Earth Science vocabularies that help ensure Earth science data and services are described in a consistent and comprehensive manner and allow for the precise searching of collection-level metadata and subsequent retrieval of data and services. Initiated over twenty years ago, the GCMD Keywords are periodically analyzed for relevancy and will continue to be refined and expanded in response to user needs. This talk explores the current status of the GCMD keywords, the value and usage that the keywords bring to different tools/agencies as it relates to data discovery, and how the keywords relate to SWEET (Semantic Web for Earth and Environmental Terminology) Ontologies.

data discover

NASA GIBS and Worldview: Bringing 20 Years of Terra Data into View

For nearly 10 years, the NASA Global Imagery Browse Services (GIBS) and Worldview interactive mapping site have provided users full resolution visualizations of Terra land, cryosphere, ocean, and atmosphere science parameters available within hours of acquisition. Over that time, GIBS and Worldview have expanded their Terra visualization suite to include even more parameters covering the entire mission of science data. Users can now view daily visualizations of nearly 50 near-real time and over 125 science quality science parameters from Terra instruments. These visualizations provide a unique capability for a broad user base to interact and discover the wealth of information sense by instruments on the Terra platform. By viewing Terra-based visualizations in a single location, users can correlate retrievals across instruments. Additionally, GIBS and Worldview include visualizations from many other platforms, allowing for an even greater data discovery capability.The GIBS and Worldview teams continue to work with Terra instrument science teams to add more visualized layers, as well as supporting visualization updates as the data is continually improved. Additionally, the GIBS interfaces and Worldview site are actively working on improved visualization functionality, including vector- and granule/swath-based products, to better address the expanding needs of its user community. This presentation will focus on an overview of existing and future visualization products and capabilities that have supported, and will support, a broad use of Terra data in science, media, and application communities.

Cechini, Matthew

Exploiting Dark Information Resources to Create New Value Added Services to Study Earth Science Phenomena

This paper presents two research applications exploiting unused metadata resources in novel ways to aid data discovery and exploration capabilities. The results based on the experiments are encouraging and each application has the potential to serve as a useful standalone component or service in a data system. There were also some interesting lessons learned while designing the two applications and these are presented next.

Earth Science Informatics

NASA Giovanni: Analyze, Compare, and Visualize 2000+ Earth Satellite and Model Variables Without Downloading Data and Software

Over vast oceans and remote continents, observations are often scarce and discontinuous. Satellite and model data play a critical role in research and applications. However, finding and accessing satellite and model data can be a daunting task for many, especially those outside the community. The NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), one of 12 NASA Science Mission Directorate Data Centers, provides Earth science data, information, and services to everyone such as researchers, application users, educators, and students. GES DISC archives and supports datasets applicable to several NASA Earth Science Focus Areas including Atmospheric Composition, Water & Energy Cycles, Carbon Cycle & Ecosystem, and Climate Variability. To facilitate data discovery, evaluation, and exploration, GES DISC has developed the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni), an online tool to analyze and visualize NASA remote sensing and model data without downloading data and software. As of this writing, over 2000 Earth satellite and model variables are available in Giovanni, including several wellknown NASA satellite missions (e.g., TRMM, GPM) and projects (e.g., MERRA-2, GPCP). Giovanni provides twenty-two plots that can be used to analyze, compare, and explore Earth data across different disciplines. Results can be shared with colleagues and downloaded for further analysis. Over the years, Giovanni has helped publish over 3000 referral papers. In this presentation, we will showcase key variables and plot types in Giovanni with examples. In particular, we will present several popular precipitation products from GPM and CPCP for evaluation and comparison.

data analysis

Managing and Servicing Physical Oceanographic Data at a NASA Distributed Active Archive Center

The NASA Earth Science Data Information Systems Project funds and operates 12 Distributed Active Archive Center(s) (DAAC) throughout the United States. Of these 12 centers, the Physical Oceanography DAAC (PO.DAAC) is committed to providing long term archival, distribution and stewardship for NASA physical oceanographic data, primarily derived from space-born satellite systems, but also including a growing set of recent and future in situ observations from the SPURS-1 and SPURS-2 campaigns. Notable NASA missions supported include: Seasat, TOPEX/Poseidon, NSCAT, QuikSCAT, ISS-RapidScat, Jason-1, Jason-2/OSTM, GRACE, Aquarius, GHRSST, and MODIS. The following interagency and international missions are also supported by PO.DAAC: AVHRR, Coriolis, DMSP, MetOp-A, MetOp-B, Oceansat-2. The PO.DAAC currently holds 525 datasets in public distribution, spanning the following observational parameters: sea surface temperature, sea surface salinity, ocean color, ocean surface currents, ocean surface wind speed, ocean surface wind direction, sea surface height, significant wave height, ocean water mass/thickness, and sea ice age. A hundred of these datasets are available in near-real-time. Datasets are distributed through a variety of open-source access protocols including FTP, OPeNDAP, and THREDDS. FTP will soon be phased out in favor of a recently introduced HTTPS PO.DAAC Drive interface that supports WebDAV and interoperable machine-to-machine communication. OPeNDAP supports remote data/metadata query, subset, and download. THREDDS provides the features of OPeNDAP with the additional feature of temporal aggregation. PO.DAAC also offers proprietary tools and services to further enhance the data discovery, visualization and analysis experience, including but not limited to: State of the Ocean, Web Services (data/metadata discovery and extraction), HiTIDE Level-2 subsetter, Live Access Server (LAS), Webification (w10nsci), and Rich Site Summary (RSS) Datacasting. To assist with provenance of datasets, PO.DAAC has implemented DOIs for the data it distributes so that they can be properly cited. There is a user forum and helpdesk that contains data recipes and via which users can get guidance. In summary, this presentation aims to provide a general overview of PO.DAAC’s web portal and data holdings along with a set of illustrative examples leading prospective data users into the practical utility of its tools and services.

Moroni, David F.

Long-term measurements of ice nucleating particles at Atmospheric Radiation Measurement (ARM) sites worldwide

Ice nucleating particles (INPs) play a critical role in cloud microphysics and precipitation formation, yet long-term, spatially extensive observational datasets remain limited. Here, we present one of the most comprehensive publicly available datasets of immersion-mode INP concentrations using a single analytical method, generated through the U.S. Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) user facility. INP filter samples have been collected across a broad range of environments – including agricultural plains, Arctic coastlines, high-elevation mountain sites, marine regions, and urban areas – via fixed observatories, mobile facility deployments, and vertically-resolved tethered balloon system operations. We describe the standardized processing and quality assurance pipeline, from filter collection and processing using the Ice Nucleation Spectrometer to final data products archived on the ARM Data Discovery portal. The dataset includes both total INP concentrations and selectively treated samples, allowing for classification of biological, organic, and inorganic INP types. It features a continuous 5-year record of INP measurements from a central U.S. site, with data collection still ongoing. Seasonal and site-specific differences in INP concentrations are illustrated through intercomparisons at −10 and −20 °C, revealing distinct regional sources and atmospheric drivers. We also outline mechanisms for researchers to access existing data, request additional sample analyses, and propose future field campaigns involving ARM INP measurements. This dataset supports a wide range of scientific applications, from observational and mechanistic studies to model development, and provides critical constraints on aerosol-cloud interactions across diverse atmospheric regimes (Creamean et al., 2024, 2020b; https://doi.org/10.5439/1770816).

Creamean, Jessie M. [Colorado State Univ., Fort Co

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew

NASA R&M Efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) Digital Assets

At the intersection of mission, technology, and place is NASA’s need to modernize for a digital-forward future. Digitalization, the process of moving toward digital business, is occurring everywhere and remains an ongoing process across the federal government.”[1] Whereas, Digital Transformation is “employing digitization/digital technologies (e.g., Artificial Intelligence (AI), mobile, cloud, data) to change a process, product, or capability so dramatically (e.g., real-time, intelligent, personalized, anywhere, anytime) that it is unrecognizable compared to its traditional form.” [2] In order to facilitate a digital transformation it is essential for NASA to understand and identify where data exists today and which data are value-needed in the future, understand where there are unfulfilled data needs that limit the advancement of NASA work, and ensure NASA efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) digital assets in the future. Therefore, NASA’s Reliability & Maintainability (R&M) Enterprise Data Sharing team is working to leverage both Digitization and Digital Transformation to achieve their vision of developing an R&M data discovery framework that enables our community, our partners, and our stakeholders with the ability to efficiently, robustly, and seamlessly access information that enables real-time knowledge and model-based, analytics driven, decision-making impacting R&M. As a result the R&M Enterprise Data Sharing team has conducted a survey of its Reliability, Maintainability, and Availability (RMA) community members to identify data existence (created or used) and where there are corresponding barriers to data acquisition and/or R&M or other issues as shown within this presentation.

Digital Transformation

Trilateral Task Force – Reliability Analysis Supporting Mission Extension/Post Mission Disposal

At the intersection of mission, technology, and place is NASA’s need to modernize for a digital-forward future. Digitalization, the process of moving toward digital business, is occurring everywhere and remains an ongoing process across the federal government.”[1] Whereas, Digital Transformation is “employing digitization/digital technologies (e.g., Artificial Intelligence (AI), mobile, cloud, data) to change a process, product, or capability so dramatically (e.g., real-time, intelligent, personalized, anywhere, anytime) that it is unrecognizable compared to its traditional form.” [2] In order to facilitate a digital transformation it is essential for NASA to understand and identify where data exists today and which data are value-needed in the future, understand where there are unfulfilled data needs that limit the advancement of NASA work, and ensure NASA efficiency through Findable, Accessible, Interoperable, and Reusable (FAIR) digital assets in the future. Therefore, NASA’s Reliability & Maintainability (R&M) Enterprise Data Sharing team is working to leverage both Digitization and Digital Transformation to achieve their vision of developing an R&M data discovery framework that enables our community, our partners, and our stakeholders with the ability to efficiently, robustly, and seamlessly access information that enables real-time knowledge and model-based, analytics driven, decision-making impacting R&M. As a result the R&M Enterprise Data Sharing team has conducted a survey of its Reliability, Maintainability, and Availability (RMA) community members to identify data existence (created or used) and where there are corresponding barriers to data acquisition and/or R&M or other issues as shown within this presentation.

Digital Transformation, Reliability Engineering

A Global Repository for Planet-Sized Experiments and Observations

Working across U.S. federal agencies, international agencies, and multiple worldwide data centers, and spanning seven international network organizations, the Earth System Grid Federation (ESGF) allows users to access, analyze, and visualize data using a globally federated collection of networks, computers, and software. Its architecture employs a system of geographically distributed peer nodes that are independently administered yet united by common federation protocols and application programming interfaces (APIs). The full ESGF infrastructure has now been adopted by multiple Earth science projects and allows access to petabytes of geophysical data, including the Coupled Model Intercomparison Project (CMIP) output used by the Intergovernmental Panel on Climate Change assessment reports. Data served by ESGF not only include model output (i.e., CMIP simulation runs) but also include observational data from satellites and instruments, reanalyses, and generated images. Metadata summarize basic information about the data for fast and easy data discovery.

Earth Systems Grid Federation (ESGFC)