Engineering PapersSearch

Engineering topics

Andrey Savtchenko

Publications and source records attributed to Andrey Savtchenko.

At least 19 records

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source

Assessing the Impacts of Two Averaging Methods on AIRS Level 3 Monthly Products and Multi-Year Monthly Means

The Atmospheric Infrared Sounder (AIRS) onboard NASA’s Aqua satellite provides more than 16 years of data. Its monthly gridded (Level 3) product has been widely used for climate research and applications. Since counts of successful soundings in a grid cell are used to derive monthly averages, this “Averaged By Observations (ABO)” approach effectively gives equal importance to all participating soundings within a month. It is conceivable, then, that days with more observations due to day-to-day orbit shift and regimes with better retrieval skills, will contribute disproportionately to the monthly average within a cell. Alternatively, the AIRS Level 3 monthly product can be produced through an "Averaged By Days (ABD)” approach, where the monthly mean in a grid cell is a simple average of the daily means. The effects of these averaging methods on the AIRS version 6 monthly product are assessed quantitatively using temperature and water vapor at surface and 500hPa. The ABO method results in a warmer (slightly colder) global mean temperature at surface (500hPa) and a drier global mean water vapor than ABD method. The AIRS multi-year monthly mean temperature and water vapor from both methods are also compared with the Modern-Era Retrospective analysis for Research and Applications – 2 (MERRA-2) product and evaluated with a simulation experiment, indicating the ABD method has less error and is more closely correlated with MERRA-2. In summary, the ABD method is recommended for future versions of the AIRS Level 3 monthly product and more data services supporting Level 3 aggregation are needed.

AIRS

NO2 Anomalies - Economy Attribution and Rapid Climate Response

Using principal component (PC) analysis of 16 years of monthly series of NO2 from OMI/Aura, we show that it is the third PC (PC3) from the full hierarchy of principal modes that is best coupled with the economic indicators. This coupling is positive, i.e. PC3 and economic indicators manifest positive covariance. However, the economic variability can explain only 40% of the information in PC3. Furthermore, this mode by itself explains only 3% of the total deseasonalized NO2 variability. We thus conclude that, while having an unambiguous impact, the economy can be awarded at best third order of importance in driving the deseasonalized NO2 series, i.e. after seasonal variability is removed from the series. Once we identified PC3 as the NO2 mode that is coupled with economic variability, we use this mode as an indicator and look for rapid climate adjustments to that part of NO2 variability that we are confident is coupled with the economic variability. We focus on observational data from the Atmospheric Infrared Sounder (AIRS) on board of NASA Aqua satellite, decompose series of surface skin temperature and clear-sky outgoing longwave radiances (OLR) into principal components, and identify potential impacts of NO2 PC3 on these climate variables.

NO2

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov

Publishing Variables Archived at GES DISC to Earth System Grid Federation (ESGF)

We present a straightforward and low-cost approach to publish variables archived at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) to the Earth System Grid Federation (ESGF). An ESGF publication requires a single standard-name variable aggregated over time to facilitate data inter-comparison. It also contains significant metadata to enable searching in ESGF. We look up standard names on high demand in ESGF search history, and using OPeNDAP and NcML technologies we aggregate the corresponding variables available in the GES DISC archive with augmented metadata required by CMIP6 and obs4MIPs Data Specification version 2.1. At this writing 10 variables from a standard product of the Atmospheric Infrared Sounder along with the Tech Notes are published in ESGF by NASA Center for Climate Simulation (NCCS). Users can view, analyze, and subset remotely, and download these aggregated variables via links in any ESGF node after searching. We plan to work on and publish more variables and data from different NASA missions and experiments in our archive.

Fan Fang

NO2 Anomalies - Economy Attribution and Rapid Climate Response

Using principal component (PC) analysis of 16 years of monthly series of NO2 from OMI/Aura, we show that it is the third PC (PC3) from the full hierarchy of principal modes that is best coupled with the economic indicators. This coupling is positive, i.e. PC3 and economic indicators manifest positive covariance. However, the economic variability can explain only 40%of the information in PC3. Furthermore, this mode by itself explains only 3% of the total deseasonalized NO2 variability. We thus conclude that, while having an unambiguous impact, the economy can be awarded at best third order of importance in driving the deseasonalized NO2 series, i.e. after seasonal variability is removed from the series. Once we identified PC3 as the NO2 mode that is coupled with economic variability, we use this mode as an indicator and look for rapid climate adjustments to that part of NO2 variability that we are confident is coupled with the economic variability. We focus on observational data from the Atmospheric Infrared Sounder (AIRS) on board of NASA Aqua satellite, decompose series of surface skin temperature and clear-sky outgoing longwave radiances (OLR) into principal components, and identify potential impacts of NO2 PC3 on these climate variables.

NO2

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov

Automated classification of scientific publications linked to GES DISC datasets

The data collections archived and distributedby the GES DISC NASA data center arewidely utilized for various Earth Science studies.As these collections are created, many researchworks are published regarding the collections, algorithms,validations and applications. SinceGES DISC collects these publications and providestheir citations for the users, it is helpful tocategorize them based on how they relate to the datasetsthey are associated with. Specifically,whether the publication that is linked to GES DISCdataset is using it for applicational research,or if it describes the algorithm for dataset creation,or the validation of the dataset, or providesthe general overview of the data collection. Currently,this process requires simple manuallabelling, and as such, may be possible to solve viaautomation. To approach this problem, wedeveloped machine learning classifiers to predictthe category a publication belongs to. We usedmanually labeled publications as training data forsupervised machine learning algorithms:Random Forest and Naive Bayes. We achieved classificationaccuracy that is substantially betterthan the baseline accuracy, thus greatly improvingthe efficiency of the publication internalanalysis.

Rohan Dayal

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal

Towards Automated Analytics of Research Publications

For readers of scientific publications it remains a big challenge to unambiguously relate the published research with the data used. To a substantial degree it is attributed to authors, journals, editors, and reviewers not prioritizing correct data citation, which impacts traceability, repeatability, and giving credits to published authors and their funding sources. Furthermore, uniform classification of the content of the published research is hampered by journals using journal specific topics and letting authors to assign free text keywords to their papers. We demonstrate automated analytics methods for extracting and relating datasets used and the research application areas by processing 1,300 research papers that referenced the NASA Giovanni service (but probably not the datasets in particular) as supporting their publication process. This presentation was given during the 2022 ESIP January meeting held virtually in January 2022.

Irina Gerasimov

Creating a Repository of Publication Citations for a Data Center

Tracking dataset citations in scientific publications provide multiple benefits: obtaining citation indices for quantitative evaluation of the dataset scientific impact, learning about dataset usage in applied sciences, credits to dataset creators, datasets co-citation relationships and many more.

Infometrics

Sampling of Space-Based Observations of XCO2 Associated With Biomass Burning Events

- The OCO-2 Level 3 assimilation product fills in gaps where OCO-2 does not have observations. - Although the Level 3 clearly shows the increased CO2 over the Amazon associated with the September 2017 biomass burning event, the total amount of CO2 may be underestimated because of a sampling bias. - The TCCON observations match fairly well with the OCO-2 observations and can be used to estimate the accuracy of the OCO observations and the temporal sampling bias. However, since it is a remotely sensed product, it may suffer the same sampling biases associated with biomass burning events as OCO-2. - Although aircraft data show that there is a significant variation in xCO2 with altitude, they suggest the OCO-2 fused products created by kriging may be better than the assimilated product at detecting large deviations from the average state on daily and regional scales.

Thomas J. Hearty, III

IMERG and GPCP Seasonality and Response to Climate and Weather Variability

The Integrated Multi-satellitE Retrievals for GPM (IMERG) and the The Global Precipitation Climatology Project (GPCP) are two of the most popular precipitation products. IMERG is a relatively new dataset that targets the needs primarily of the hydrological community by resolving the hydrological cycle of precipitation at fine temporal (30-minutes) and spatial (10-km) scales. IMERG only recently exceeded the user base of the highly successful, but discontinued in 2019, TRMM Multi-satellite Precipitation Analysis (TMPA). GPCP, on the other hand, has traditionally been strong in the climate research community, and recently has been revised under the framework of NASA's Making Earth System Data records for Use in Research Environments (MEaSUREs) program. Both IMERG and GPCP are similar in the underlying approaches to achieve global coverage, in particular using satellite microwave and infrared observations, and adjusting the precipitation retrieval with rain gauge information. While there is a tendency to use both datasets interchangeably, differences remain and some of them limit the usage of IMERG as a climate data record at this point. By applying Principal Component Analysis, we identify the differences and similarities mode-by-mode, and further the guidance on suggested usage of IMERG as a climate data record. As an example in the attached figure, IMERG vs GPSP differences in the explained variance by the two leading seasonal modes can be identified, mainly in the extreme southern latitudes, and over western boundary currents (Gulfstream and Kuroshio). At the preparation time for this presentation, the new version "07"of IMERG was in the works that may resolve the issues presented here. Nevertheless, our analysis can help to gauge the uncertainties of the studies already done using the currently existing IMERG version "06", and evaluate the improvements in the upcoming version "07".

Andrey Savtchenko

Utilizing Google Scholar as a Bibliographic Resource for Publications Search

This study focuses on the automated search for publication citations for the Earth Observing System Data and Information System (EOSDIS) datasets. The research investigates the feasibility of using automated search methods to gather published works from various bibliometric databases. A comparison is presented, highlighting the differences in citation counts obtained from Google Scholar compared to established bibliographic databases. The study also introduces a methodology and an open-source tool for getting publication citations from Google Scholar, utilizing dataset DOIs and keyword searches. The findings contribute to understanding the reliability and effectiveness of Google Scholar as a source for dataset citation retrieval and provide researchers with a valuable resource for obtaining comprehensive citation data.

Infometrics

Harmonizing Multi-Disciplinary Data for Applied Research: A Comprehensive Analysis of GES DISC Datasets through NLP and TF-IDF Techniques

Data centers distribute data encompassing multiple disciplines, making it necessary to evaluate dataset applicability to applied research. Typically, these datasets are generated by specialized science teams and consist of single-discipline data, such as atmospheric temperature, pressure, and precipitation, leading to unique dataset formats and access services. However, in applied research, the utilization of datasets from multiple disciplines is commonly necessary. The increasing availability of research literature citing Earth Science datasets presents an opportunity to analyze the usage of datasets in multi-disciplinary research. This study proposes a novel approach wherein research publications citing datasets archived at the GES DISC (Goddard Earth Sciences Data and Information Services Center) are collected, and each publication is associated with specific research topics through the application of Natural Language Processing (NLP), using the Term Frequency Inverse Document Frequency (TF-IDF) technique on the publication titles and abstracts. Through this analysis, we gain insights into the distribution of dataset disciplines as they are being used in various applied research areas. This knowledge is essential for the development of dataset tools and services tailored to effectively support applied research studies, as it enhances a data center's comprehension of how datasets from multiple disciplines are integrated into research endeavors.

Infometrics