Interdisciplinary Community Needs- Water-Energy-Food Nexus
Explore the source record for details and available documents.
Engineering topics
Publications and source records attributed to Irina Gerasimov.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Introduction to Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) Giovanni … is a Web-based visualization and analysis system that provides 22 different visualization and analysis options, operating on thousands of Earth science data variables generated by satellite instrument observations and from related model datasets Giovanni … was originally conceived as a data exploration tool, but its ease-of-use, analytical capabilities (spatial and temporal subsetting, multi-period averaging, data mapping and time-series, and more) have led to its use as a multi-discipline research tool Giovanni … provided unprecedented access to NASA Earth science data for many different disciplines, AND is still providing a simple way to find, analyze, visualize, and utilize such data for a wide spectrum of research topics
Explore the source record for details and available documents.
NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.
The TRopospheric Ozone and Precursors from Earth System Sounding (TROPESS) project generates Earth System Data Records (ESDRs) of ozone, and other atmospheric constituents (CH4,CO, H2O, HDO, NH3, PAN and temperature) by processing data from multiple satellites through a common retrieval algorithm and ground data system. Satellite Level-1B input data used in generating the TROPESS L2 data products include CrIS NOAA-20 (JPSS-1), CrIS SNPP, AIRS Aqua, OMI Aura, and TROPOMI S5P. The common retrieval framework is known as the MUlti-SpEctra, MUlti-SpEcies, Multi-SEnsors (MUSES) science data processing system (MUSES-SDPS). Several of the TROPESS data products are now available from the NASA Goddard Earth Sciences Data and Information Service Center (GES DISC) for users to download. In this presentation we provide an overview of the various TROPESS data products. These data products can be divided into the following Forward Stream types: Standard Products, Summary Products, and Full-Archival Products. Standard Products are for users that are doing full analysis with avenging kernel and covariance corresponding to retrieved vertical profiles. Summary products have a smaller file size and are more convenient for first-look and rapid analysis, include total and partial columns, as well as column averaging kernels. The Full-Archival Products will contain all information used in creating the data. TROPESS also creates Special Products, provided on an as-needed and as-available basis to support NASA field missions and individual-investigator requests over specific regions. Eventually, TROPESS will also produce and deliver a set of Reanalysis Stream products. Data at the GES DISC are being transitioned into the "Cloud". This will allow users with "Cloud” access to perform data analysis directly on the data without downloading the data to their system. Services, such as subsetting and data visualization, will also be provided for TROPESS data products at the GES DISC.
The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has been actively involved in many aspects of ensuring the long-term preservation of NASA earth science data and knowledge. This involves both the recovery and preservation of early NASA meteorological and other earth observation data, as well as preserving the more recent Earth Observation System (EOS) mission data sets which continue or have reached their end of lifetime. The GES DISC adds value to these preserved data by adding metadata and making the data available online to future researchers. The early NASA meteorological and earth observation data sets from the 1960s and 70s were originally archived on magnetic tapes, and visualizations of these data were preserved on 70-mm film. As these media have aged, their contents have been at risk of permanent loss. NASA has given the task of preserving these early data sets to the GES DISC and making these data sets easily available to the public. The data from these early missions are potentially useful to climate researchers as these are some of the only global measurements made at their time. These old data on magnetic tapes and film strips do not contain easily readable metadata, and so to add value the GES DISC has added digital metadata to them so that the data are searchable and findable. The GES DISC is also involved in preserving the data and knowledge from the EOS era missions. The GES DISC follows the guidelines developed for the preservation of data as specified in the NASA EOS Data and Information System (EOSDIS) Earth Science Data Preservation Content Specification (423-SPEC-001) document. To date, the GES DISC has consulted with the data science teams from the following missions: UARS, Earth Probe TOMS, Aura HIRDLS, and SORCE, in order to properly preserve their data and accompanying documentation. The GES DISC is also currently working with the EOS science teams from TRMM, AIRS, MLS, OMI and additional missions to ensure that the relevant documents and data sets are properly archived for future researchers. A standardized procedure for mission data preservation following 423-SPEC-001 makes preservation among the many NASA EOSDIS data centers uniform, so that these could be transitioned easily to a common EOSDIS preservation repository. This presentation will give an overview of the preservation and recovery of the old NASA historical data sets archived at the GES DISC, as well as the data and documentation preservation efforts of the EOS era missions.
The NASA Modern-Era Retrospective analysis for Research and Applications Version 2 (MERRA-2) is atmospheric reanalysis data spanning 1980 to present. It has been produced by the NASA Global Modeling and Assimilation Office (GMAO) and is distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). MERRA-2 data includes 100 collections of Earth system variables, mainly from the atmospheric model, such as aerosol fields and meteorological fields, radiation fields, and aerosol fields, guided by the assimilation of as many as six million observations every six hours. MERRA-2 has been one of the most popular datasets from NASA and is widely used in interdisciplinary research and applications, with increasing numbers of new users. For example, at least 7000 users accessed MERRA-2 data at GES DISC in the year 2021, ~1000 more users than in the year 2020. In this presentation, we will introduce the MERRA-2 datasets associated with aerosol and air quality studies and use a wildfire case study to demonstrate the data tools developed at GES DISC to analyze and visualize MERRA-2 data, such as Giovanni and the level 3 and level 4 subsetter, and Jupyter Python notebook. We will also update the status of cloud migration of the MERRA-2 data to Amazon Web Services (AWS).
Explore the source record for details and available documents.
This study focuses on the automated search for publication citations for the Earth Observing System Data and Information System (EOSDIS) datasets. The research investigates the feasibility of using automated search methods to gather published works from various bibliometric databases. A comparison is presented, highlighting the differences in citation counts obtained from Google Scholar compared to established bibliographic databases. The study also introduces a methodology and an open-source tool for getting publication citations from Google Scholar, utilizing dataset DOIs and keyword searches. The findings contribute to understanding the reliability and effectiveness of Google Scholar as a source for dataset citation retrieval and provide researchers with a valuable resource for obtaining comprehensive citation data.
Data centers distribute data encompassing multiple disciplines, making it necessary to evaluate dataset applicability to applied research. Typically, these datasets are generated by specialized science teams and consist of single-discipline data, such as atmospheric temperature, pressure, and precipitation, leading to unique dataset formats and access services. However, in applied research, the utilization of datasets from multiple disciplines is commonly necessary. The increasing availability of research literature citing Earth Science datasets presents an opportunity to analyze the usage of datasets in multi-disciplinary research. This study proposes a novel approach wherein research publications citing datasets archived at the GES DISC (Goddard Earth Sciences Data and Information Services Center) are collected, and each publication is associated with specific research topics through the application of Natural Language Processing (NLP), using the Term Frequency Inverse Document Frequency (TF-IDF) technique on the publication titles and abstracts. Through this analysis, we gain insights into the distribution of dataset disciplines as they are being used in various applied research areas. This knowledge is essential for the development of dataset tools and services tailored to effectively support applied research studies, as it enhances a data center's comprehension of how datasets from multiple disciplines are integrated into research endeavors.
Explore the source record for details and available documents.
● In the evolving landscape of open science, the ability to navigate and discover pertinent datasets is increasingly significant. This primarily hinges on the presence of detailed metadata, delineating the dataset’s content, and potential spheres of application. ● The GES DISC datasets are characterized by science keywords to enable dataset discovery in web search interfaces. ● A problem may arise where a dataset lacks a science keyword that it otherwise should have. ● Machine learning techniques such as link prediction can be used to detect these missing science keywords by estimating the probability of new links forming between dataset and keyword nodes.
Explore the source record for details and available documents.
The NASA GES DISC is the primary archive for the Aura OMI products. After 20 years of operations, the OMI instrument is nearing its end and entering Phase-F (closeout). It is time to begin the process of preserving the data, documentation, software and associated information that were produced by the OMI project over its lifetime.
In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences." "In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences.
Scientific datasets are increasingly cited in peer-reviewed journal publications, facilitating easy access to research utilizing those datasets. Datasets undergo a life cycle where older versions of datasets are replaced by newer versions often due to improvements in data resolution, algorithms, and other factors. Unlike peer reviewed documents registered with a single Digital Unique Identifier (DOI), datasets can be updated over time and the newer version of the datasets are registered with a new DOI which is not necessarily linked to the previous version of the dataset. It is challenging when publications citing a dataset need to be traced over the entire life cycle of that dataset. We provide an innovative approach to link the dataset versions and publications using a knowledge graph (KG). KG can help to trace the dataset cited in publications over the entire dataset life cycle and shed light into dataset usage in various applied research areas. We fine-tuned the pretrained NASA IMPACTINDUS Large Language Model (LLM) on a set of labeled publications abstracts. Our results showed that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were classified into non-applied research areas. By linking datasets to applied research areas through the KG and employing Global Change Master Directory(GCMD), a well-established controlled vocabulary of scientific keywords describing Earth science datasets, we contribute to a transparent and advanced search and discovery mechanism for datasets across the Earth data ecosystem. The integrated KG and LLM approach is now incorporated and operational in dataset publication management at one of NASA’s Earth science data archival centers.