Engineering PapersSearch

Engineering topics

Jerome Alfred

Publications and source records attributed to Jerome Alfred.

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov

Towards Automated Analytics of Research Publications

For readers of scientific publications it remains a big challenge to unambiguously relate the published research with the data used. To a substantial degree it is attributed to authors, journals, editors, and reviewers not prioritizing correct data citation, which impacts traceability, repeatability, and giving credits to published authors and their funding sources. Furthermore, uniform classification of the content of the published research is hampered by journals using journal specific topics and letting authors to assign free text keywords to their papers. We demonstrate automated analytics methods for extracting and relating datasets used and the research application areas by processing 1,300 research papers that referenced the NASA Giovanni service (but probably not the datasets in particular) as supporting their publication process. This presentation was given during the 2022 ESIP January meeting held virtually in January 2022.

Irina Gerasimov

Creating a Repository of Publication Citations for a Data Center

Tracking dataset citations in scientific publications provide multiple benefits: obtaining citation indices for quantitative evaluation of the dataset scientific impact, learning about dataset usage in applied sciences, credits to dataset creators, datasets co-citation relationships and many more.

Infometrics

Automated Collection of Scientific Publications Linked to NASA Earth Science Datasets

NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.

Infometrics

Utilizing Google Scholar as a Bibliographic Resource for Publications Search

This study focuses on the automated search for publication citations for the Earth Observing System Data and Information System (EOSDIS) datasets. The research investigates the feasibility of using automated search methods to gather published works from various bibliometric databases. A comparison is presented, highlighting the differences in citation counts obtained from Google Scholar compared to established bibliographic databases. The study also introduces a methodology and an open-source tool for getting publication citations from Google Scholar, utilizing dataset DOIs and keyword searches. The findings contribute to understanding the reliability and effectiveness of Google Scholar as a source for dataset citation retrieval and provide researchers with a valuable resource for obtaining comprehensive citation data.

Infometrics

Harmonizing Multi-Disciplinary Data for Applied Research: A Comprehensive Analysis of GES DISC Datasets through NLP and TF-IDF Techniques

Data centers distribute data encompassing multiple disciplines, making it necessary to evaluate dataset applicability to applied research. Typically, these datasets are generated by specialized science teams and consist of single-discipline data, such as atmospheric temperature, pressure, and precipitation, leading to unique dataset formats and access services. However, in applied research, the utilization of datasets from multiple disciplines is commonly necessary. The increasing availability of research literature citing Earth Science datasets presents an opportunity to analyze the usage of datasets in multi-disciplinary research. This study proposes a novel approach wherein research publications citing datasets archived at the GES DISC (Goddard Earth Sciences Data and Information Services Center) are collected, and each publication is associated with specific research topics through the application of Natural Language Processing (NLP), using the Term Frequency Inverse Document Frequency (TF-IDF) technique on the publication titles and abstracts. Through this analysis, we gain insights into the distribution of dataset disciplines as they are being used in various applied research areas. This knowledge is essential for the development of dataset tools and services tailored to effectively support applied research studies, as it enhances a data center's comprehension of how datasets from multiple disciplines are integrated into research endeavors.

Infometrics

Discovering Research Areas in Dataset Applications Through Knowledge Graphs and Large Language Models

Scientific datasets are increasingly cited in peer-reviewed journal publications, facilitating easy access to research utilizing those datasets. Datasets undergo a life cycle where older versions of datasets are replaced by newer versions often due to improvements in data resolution, algorithms, and other factors. Unlike peer reviewed documents registered with a single Digital Unique Identifier (DOI), datasets can be updated over time and the newer version of the datasets are registered with a new DOI which is not necessarily linked to the previous version of the dataset. It is challenging when publications citing a dataset need to be traced over the entire life cycle of that dataset. We provide an innovative approach to link the dataset versions and publications using a knowledge graph (KG). KG can help to trace the dataset cited in publications over the entire dataset life cycle and shed light into dataset usage in various applied research areas. We fine-tuned the pretrained NASA IMPACTINDUS Large Language Model (LLM) on a set of labeled publications abstracts. Our results showed that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were classified into non-applied research areas. By linking datasets to applied research areas through the KG and employing Global Change Master Directory(GCMD), a well-established controlled vocabulary of scientific keywords describing Earth science datasets, we contribute to a transparent and advanced search and discovery mechanism for datasets across the Earth data ecosystem. The integrated KG and LLM approach is now incorporated and operational in dataset publication management at one of NASA’s Earth science data archival centers.

data provenance

Enhancing Dataset Discovery and Usage Tracking in Earth Sciences: Integrating Knowledge Graphs and Large Language Models

NASA's Data Active Archive Centers (DAACs) have played a crucial role in supporting a wide range of applied research in Earth and Environmental sciences. To date, over 20,000 publications have been collected, citing more than 3,000 NASA Earth science datasets. We present an innovative approach that links datasets and collected publications through a knowledge graph (KG). This KG enables the tracking of dataset citations throughout the dataset's lifecycle, revealing patterns of dataset usage across various applied research areas. We fine-tuned the pre-trained NASA IMPACT INDUS-Base Retriever Large Language Model (LLM) using a set of labeled publication abstracts. Our results indicate that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were categorized into non-applied research areas. The classified publications linked to datasets are used to discover datasets by users interested in specific applied research and by dataset providers to determine dataset usage for applications.

open-source