Engineering PapersSearch

Engineering topics

Binita KC

Publications and source records attributed to Binita KC.

At least 19 records

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source

Updates of MERRA-2 Data and Services at NASA GES DISC

Over40 years of NASA climate reanalysis datasets from the Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) are available at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). In addition to being used in traditional weather and climate research, MERRA-2 is also widely used in application studies of, e.g., wind and solar energy, air quality and health, food and drought, and heat waves. Two new MERRA-2 datasets were recently added at the GES DISC: (1) climate statistics derived fromMERRA-2 daily data to assist in the analysis of extreme temperature and precipitation events and of large-scale meteorological patterns from 1980 to the present and (2) gridded satellite and conventional observations processed in the MERRA-2 system, along with key statistics derived from the data assimilation, to help better understand how the quality of observations directly affect there analysis data. The GES DISC focuses its efforts on continually improving existing data services and to develop new data tools to satisfy various user communities. The newly added features include the following: Time series service: This is a new service for MERRA-2 data, which enables the easy and fast access of long-term hourly or daily time series at a location for popular parameters. The data is saved in a single le in Ascii format with a user- friendly structure. New analytic functions in the subsetter interface: Options for downloading daily minimum and maximum values have been added into the subsetter interface, in addition to the existing daily mean option, for all MERRA-2 and MERRA sub-daily products. Data format conversion to GeoTIFF has been implemented. New variables in Giovanni: Most monthly variables have been integrated into Giovanni, GESDISC’s online visualization and analysis tool. Due to the large data volume, hourly variables were selected based on user requests. More online information: New MERRA-2 documentation has been added: Data How-To, Data in Action, and FAQ. This presentation overviews two new MERRA-2 datasets and illustrates the new features of data services through a number of case studies. MERRA-2 data and services can be found at: https://disc.gsfc.gov/datasets?

Data management

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian

Towards Automated Analytics of Research Publications

For readers of scientific publications it remains a big challenge to unambiguously relate the published research with the data used. To a substantial degree it is attributed to authors, journals, editors, and reviewers not prioritizing correct data citation, which impacts traceability, repeatability, and giving credits to published authors and their funding sources. Furthermore, uniform classification of the content of the published research is hampered by journals using journal specific topics and letting authors to assign free text keywords to their papers. We demonstrate automated analytics methods for extracting and relating datasets used and the research application areas by processing 1,300 research papers that referenced the NASA Giovanni service (but probably not the datasets in particular) as supporting their publication process. This presentation was given during the 2022 ESIP January meeting held virtually in January 2022.

Irina Gerasimov

A Newly Developing Community-Oriented Data System from NASA GES DISC

Data services are essential to facilitate data access and to aid efficiency of conducting research and application activities. With emerging technologies such as cloud computing and AI/ML (Artificial Intelligence/Machine Learning) leading the pace of the data world, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), home to the permanent archive for multidisciplinary Earth Observation (EO) geospatial data to study atmospheric composition, weather and climate variability, and water and energy cycles is no exception.Interfacing directly with users as part of data center work, we understand the challenges for the required time and effort to discover, visualize, and analyze large varieties and quantities of Earth Observation information for research, monitoring, and decision-making, largely due to the existing data and information systems aim to support experienced users, but has been proved difficult for non-earth scientists and new users that are unfamiliar with the variety of formats and structures in which data, metadata, and information are stored, as well as the required methods to use them. To address these challenges, I will update our latest activities with regard to water-and energy-related products and community-oriented and user-friendly services at the GES DISC, including our plans for the emerging technologies.

Jennifer Wei

Resources at GES DISC for Agriculture Studies

● Introduction to GES DISC (Distributed Active Archive Center- DAAC) ○ Data Holding for Agriculture study ● Resources for Agricultural studies at GES DISC ○ Data Access ○ Data Services and Tools ○ Learning resources : ■ Studying seasonality ■ Studying extreme events: Floods, Flash Floods, Heat Waves, Drought ■ Data Recipe (How-To) ● Demo and hands-on

Zhong Liu

Automated Collection of Scientific Publications Linked to NASA Earth Science Datasets

NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.

Infometrics

The Cloud: Obstacles and Barrier Encountered By Users

Increasing exposure to and adoption of the Earthdata Cloud yields new concerns from users include: financial resource allotment, steep learning curves, institutional support, and cloud-readiness of data. We present common concerns expressed by early cloud adopters, and encourage conference-goers to relay their own or users's experiences. We also encourage brainstorming for how open science principles can help solve some of the barriers and obstacles users encounter with the cloud.

Alexis Hunzinger

How GES DISC is Opening Doors to Open Science

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is one of 12 NASA Earth Observing System data centers that document, process, archive, and distribute data from Earth science missions and related projects. GES DISC hosts numerous Earth Science products through various data formats and supports their access with several programming and website tools. The User Needs team at GES DISC is responsible for producing tutorials and resources to teach these tools, collecting and engaging with user feedback, and collaborating with other NASA data centers to enable science creation for everyone. This poster presentation outlines how GES DISC, and the User Needs team, are committed to enabling open science.

Christopher Battisto

Utilizing Google Scholar as a Bibliographic Resource for Publications Search

This study focuses on the automated search for publication citations for the Earth Observing System Data and Information System (EOSDIS) datasets. The research investigates the feasibility of using automated search methods to gather published works from various bibliometric databases. A comparison is presented, highlighting the differences in citation counts obtained from Google Scholar compared to established bibliographic databases. The study also introduces a methodology and an open-source tool for getting publication citations from Google Scholar, utilizing dataset DOIs and keyword searches. The findings contribute to understanding the reliability and effectiveness of Google Scholar as a source for dataset citation retrieval and provide researchers with a valuable resource for obtaining comprehensive citation data.

Infometrics