Engineering PapersSearch

SEARCH · Engineering Papers

Results for “dataset DOI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Discovering Research Areas in Dataset Applications Through Knowledge Graphs and Large Language Models

Scientific datasets are increasingly cited in peer-reviewed journal publications, facilitating easy access to research utilizing those datasets. Datasets undergo a life cycle where older versions of datasets are replaced by newer versions often due to improvements in data resolution, algorithms, and other factors. Unlike peer reviewed documents registered with a single Digital Unique Identifier (DOI), datasets can be updated over time and the newer version of the datasets are registered with a new DOI which is not necessarily linked to the previous version of the dataset. It is challenging when publications citing a dataset need to be traced over the entire life cycle of that dataset. We provide an innovative approach to link the dataset versions and publications using a knowledge graph (KG). KG can help to trace the dataset cited in publications over the entire dataset life cycle and shed light into dataset usage in various applied research areas. We fine-tuned the pretrained NASA IMPACTINDUS Large Language Model (LLM) on a set of labeled publications abstracts. Our results showed that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were classified into non-applied research areas. By linking datasets to applied research areas through the KG and employing Global Change Master Directory(GCMD), a well-established controlled vocabulary of scientific keywords describing Earth science datasets, we contribute to a transparent and advanced search and discovery mechanism for datasets across the Earth data ecosystem. The integrated KG and LLM approach is now incorporated and operational in dataset publication management at one of NASA’s Earth science data archival centers.

data provenance

Utilizing Google Scholar as a Bibliographic Resource for Publications Search

This study focuses on the automated search for publication citations for the Earth Observing System Data and Information System (EOSDIS) datasets. The research investigates the feasibility of using automated search methods to gather published works from various bibliometric databases. A comparison is presented, highlighting the differences in citation counts obtained from Google Scholar compared to established bibliographic databases. The study also introduces a methodology and an open-source tool for getting publication citations from Google Scholar, utilizing dataset DOIs and keyword searches. The findings contribute to understanding the reliability and effectiveness of Google Scholar as a source for dataset citation retrieval and provide researchers with a valuable resource for obtaining comprehensive citation data.

Infometrics

Automated Collection of Scientific Publications Linked to NASA Earth Science Datasets

NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.

Infometrics

Enhancing Dataset Discovery and Usage Tracking in Earth Sciences: Integrating Knowledge Graphs and Large Language Models

NASA's Data Active Archive Centers (DAACs) have played a crucial role in supporting a wide range of applied research in Earth and Environmental sciences. To date, over 20,000 publications have been collected, citing more than 3,000 NASA Earth science datasets. We present an innovative approach that links datasets and collected publications through a knowledge graph (KG). This KG enables the tracking of dataset citations throughout the dataset's lifecycle, revealing patterns of dataset usage across various applied research areas. We fine-tuned the pre-trained NASA IMPACT INDUS-Base Retriever Large Language Model (LLM) using a set of labeled publication abstracts. Our results indicate that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were categorized into non-applied research areas. The classified publications linked to datasets are used to discover datasets by users interested in specific applied research and by dataset providers to determine dataset usage for applications.

open-source

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov

AgMIP-Wheat Multi-Model Simulations on Climate Change Impact and Adaptation for Global Wheat

The climate change impact and adaptation simulations from the Agricultural Model Intercomparison and Improvement Project (AgMIP) for wheat provide a unique dataset of multi-model ensemble simulations for 60 representative global locations covering all global wheat mega environments. The multi-model ensemble reported here has been thoroughly benchmarked against a large number of experimental data, including different locations, growing season temperatures, atmospheric CO2 concentration, heat stress scenarios, and their interactions. In this paper, we describe the main characteristics of this global simulation dataset. Detailed cultivar, crop management, and soil datasets were compiled for all locations to drive 32 wheat growth models. The dataset consists of 30-year simulated data including 25 output variables for nine climate scenarios, including Baseline (1980-2010) with 360 or 550 ppm CO2, Baseline +2oC or +4oC with 360 or 550 ppm CO2, a mid-century climate change scenario (RCP8.5, 571 ppm CO2), and 1.5°C (423 ppm CO2) and 2.0oC (487 ppm CO2) warming above the pre-industrial period (HAPPI). This global simulation dataset can be used as a benchmark from a well-tested multi-model ensemble in future analyses of global wheat. Also, resource use efficiency (e.g., for radiation, water, and nitrogen use) and uncertainty analyses under different climate scenarios can be explored at different scales. The DOI for the dataset is 10.5281/zenodo.4027033 (AgMIP-Wheat, 2020), and all the data are available on the data repository of Zenodo (http://doi.org/10.5281/zenodo.4027033). Two scientific publications have been published based on some of these data here.

Agricultural Model Intercomparison and Improvement

NASA's Big Earth Data Initiative Accomplishments

The goal of NASA's effort for BEDI is to improve the usability, discoverability, and accessibility of Earth Observation data in support of societal benefit areas. Accomplishments: In support of BEDI goals, datasets have been entered into Common Metadata Repository(CMR), made available via the Open-source Project for a Network Data Access Protocol (OPeNDAP), have a Digital Object Identifier (DOI) registered for the dataset, and to support fast visualization many layers have been added in to the Global Imagery Browse Services (GIBS).

CMR

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source

Managing and Servicing Physical Oceanographic Data at a NASA Distributed Active Archive Center

The NASA Earth Science Data Information Systems Project funds and operates 12 Distributed Active Archive Center(s) (DAAC) throughout the United States. Of these 12 centers, the Physical Oceanography DAAC (PO.DAAC) is committed to providing long term archival, distribution and stewardship for NASA physical oceanographic data, primarily derived from space-born satellite systems, but also including a growing set of recent and future in situ observations from the SPURS-1 and SPURS-2 campaigns. Notable NASA missions supported include: Seasat, TOPEX/Poseidon, NSCAT, QuikSCAT, ISS-RapidScat, Jason-1, Jason-2/OSTM, GRACE, Aquarius, GHRSST, and MODIS. The following interagency and international missions are also supported by PO.DAAC: AVHRR, Coriolis, DMSP, MetOp-A, MetOp-B, Oceansat-2. The PO.DAAC currently holds 525 datasets in public distribution, spanning the following observational parameters: sea surface temperature, sea surface salinity, ocean color, ocean surface currents, ocean surface wind speed, ocean surface wind direction, sea surface height, significant wave height, ocean water mass/thickness, and sea ice age. A hundred of these datasets are available in near-real-time. Datasets are distributed through a variety of open-source access protocols including FTP, OPeNDAP, and THREDDS. FTP will soon be phased out in favor of a recently introduced HTTPS PO.DAAC Drive interface that supports WebDAV and interoperable machine-to-machine communication. OPeNDAP supports remote data/metadata query, subset, and download. THREDDS provides the features of OPeNDAP with the additional feature of temporal aggregation. PO.DAAC also offers proprietary tools and services to further enhance the data discovery, visualization and analysis experience, including but not limited to: State of the Ocean, Web Services (data/metadata discovery and extraction), HiTIDE Level-2 subsetter, Live Access Server (LAS), Webification (w10nsci), and Rich Site Summary (RSS) Datacasting. To assist with provenance of datasets, PO.DAAC has implemented DOIs for the data it distributes so that they can be properly cited. There is a user forum and helpdesk that contains data recipes and via which users can get guidance. In summary, this presentation aims to provide a general overview of PO.DAAC’s web portal and data holdings along with a set of illustrative examples leading prospective data users into the practical utility of its tools and services.

Moroni, David F.

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins

Regional variability of dust single scattering albedo due to mineral composition

Nearly all Earth System Models (ESMs) assume globally homogeneous dust aerosols, neglecting regional variations of the imaginary refractive index (IRI) due to varying mineral composition. This has led to a range of single scattering albedo (SSA) and direct radiative forcing (DRF) estimates, as models assign global properties using dust measurements from different regions. We use model and observational data to assess to what extent regionally varying mineral composition affects visible-band SSA and short-wave DRF of dust. We run global simulations with NASA GISS ModelE2.1, using optical properties for minerals based on laboratory-derived empirical relationships between dust IRI at visible wavelengths and the mineral content of iron oxides. When allowing mineral variations instead of homogeneous dust, we find regional differences in dust SSA up to ~0.06 and consequent variations in DRF up to ~6 and ~5 W/m2, at surface and top-of-atmosphere respectively. We compare model SSA with AERONET inversion data (Version 3.0, Level 2), filtered by aerosol size and optical properties to identify pure dust scenes. To investigate possible contamination by biomass burning aerosols, we use the dataset of Schuster et al. (2016, doi:10.5194/acp-16-1565-2016), who separately calculated the contribution of iron oxides and carbonaceous species to AERONET extinction and absorption optical depths. Our results show that: (1) the range of observed SSA (even without carbonaceous species) is larger than that resulting from homogeneous dust; (2) residual amounts of fine mode black and brown carbon may affect AERONET SSA even in seasons with limited biomass burning; (3) model SSA from our mineral scheme exceeds the AERONET range, possibly due also to an uncertain treatment of goethite. In summary, homogeneous dust cannot explain the variability of observed SSA, so regionally varying optical properties based on mineral content are necessary. Higher accuracy in soil mineralogy maps is needed to better reproduce SSA variations at regional scales. Also, distinct soil maps for hematite and goethite are required, given the high sensitivity of model SSA to their extreme optical properties.

V. Obiso

Stewardship of NASA's Earth Science Data and Ensuring Long-Term Active Archives

Program, NASA has followed an open data policy, with non-discriminatory access to data with no period of exclusive access. NASA has well-established processes for assigning and or accepting datasets into one of 12 Distributed Active Archive Centers (DAACs) that are parts of EOSDIS. EOSDIS has been evolving through several information technology cycles, adapting to hardware and software changes in the commercial sector. NASA is responsible for maintaining Earth science data as long as users are interested in using them for research and applications, which is well beyond the life of the data gathering missions. For science data to remain useful over long periods of time, steps must be taken to preserve: (1) Data bits with no corruption, (2) Discoverability and access, (3) Readability, (4) Understandability, (5) Usability' and (6). Reproducibility of results. NASAs Earth Science data and Information System (ESDIS) Project, along with the 12 EOSDIS Distributed Active Archive Centers (DAACs), has made significant progress in each of these areas over the last decade, and continues to evolve its active archive capabilities. Particular attention is being paid in recent years to ensure that the datasets are published in an easily accessible and citable manner through a unified metadata model, a common metadata repository (CMR), a coherent view through the earthdata.gov website, and assignment of Digital Object Identifiers (DOI) with well-designed landing product information pages.

Data Management

Persistent Identifiers Implementation in EOSDIS

This presentation provides the motivation for and status of implementation of persistent identifiers in NASA's Earth Observation System Data and Information System (EOSDIS). The motivation is provided from the point of view of long-term preservation of datasets such that a number of questions raised by current and future users can be answered easily and precisely. A number of artifacts need to be preserved along with datasets to make this possible, especially when the authors of datasets are no longer available to address users questions. The artifacts and datasets need to be uniquely and persistently identified and linked with each other for full traceability, understandability and scientific reproducibility. Current work in the Earth Science Data and Information System (ESDIS) Project and the Distributed Active Archive Centers (DAACs) in assigning Digital Object Identifiers (DOI) is discussed as well as challenges that remain to be addressed in the future.

persistent identifiers

Robustness of Vegetation Optical Depth Retrievals Based on L-Band Global Radiometry

Microwave vegetation optical depth (VOD) and soil moisture (SM) can be simultaneously retrieved based on L-band radiometry with polarization information. VOD is indicative of the vegetation water content (VWC) because it captures the extinction of land surface emission. If the connectivity of VOD to VWC is robust, the pair of VWC-SM observations can be viable bases for understanding soil–plant–atmosphere water relations, providing new perspectives on ecosystem science. Simultaneous SM–VOD retrievals are feasible by inverting the τ−ω model with two independent datasets in dual-channel algorithms. However, given correlated satellite vertical and horizontal brightness temperatures (TBs; TB v and TB h ), an ill-posed inverse problem arises where TB errors result in high uncertainties of retrievals. In this study, we apply the degrees-of-information (DoI) metric and propose a signal-to-noise ratio (SNR) metric to assess the “retrievability” of VOD given the Soil Moisture Active Passive (SMAP) TB v –TB h linear dependence. The application of these metrics allows determining where the VOD retrievals are robust and reliable. This is a necessary step in supporting the applications of VOD in ecology and hydrology. Results show that regions with mainly nonwoody vegetation have the best potential for VOD retrievals, though regularization is necessary. We then assess VOD time variations from two regularization products that reduce the impact of underdetermined inversions: the L3 dual-channel algorithm (L3-DCA) and the multitemporal dual-channel algorithm (MTDCA), which constrain VOD time dynamics with and without using a priori VOD climatology, respectively. Though they both reduce noise, especially in the VOD retrievals, they result in differences in VOD seasonal amplitude and coupling to SM at high frequencies as we outline here.

Microwave

FLUXNET-CH4: a global, multi-ecosystem dataset and analysis of methane seasonality from freshwater wetlands

Methane (CH4) emissions from natural landscapes constitute roughly half of global CH4 contributions to the atmosphere, yet large uncertainties remain in the absolute magnitude and the seasonality of emission quantities and drivers. Eddy covariance (EC) measurements of CH4 flux are ideal for constraining ecosystem-scale CH4 emissions due to quasi-continuous and high-temporal-resolution CH4 flux measurements, coincident carbon dioxide, water, and energy flux measurements, lack of ecosystem disturbance, and increased availability of datasets over the last decade. Here, we (1) describe the newly published dataset, FLUXNET-CH4 Version 1.0, the first open-source global dataset of CH4 EC measurements (available at https://fluxnet.org/data/fluxnet-ch4-community-product/, last access: 7 April 2021). FLUXNET-CH4 includes half-hourly and daily gap-filled and non-gap-filled aggregated CH4 fluxes and meteorological data from 79 sites globally: 42 freshwater wetlands, 6 brackish and saline wetlands, 7 formerly drained ecosystems, 7 rice paddy sites, 2 lakes, and 15 uplands. Then, we (2) evaluate FLUXNET-CH4 representativeness for freshwater wetland coverage globally because the majority of sites in FLUXNET-CH4 Version 1.0 are freshwater wetlands which are a substantial source of total atmospheric CH4 emissions; and (3) we provide the first global estimates of the seasonal variability and seasonality predictors of freshwater wetland CH4 fluxes. Our representativeness analysis suggests that the freshwater wetland sites in the dataset cover global wetland bioclimatic attributes (encompassing energy, moisture, and vegetation-related parameters) in arctic, boreal, and temperate regions but only sparsely cover humid tropical regions. Seasonality metrics of wetland CH4 emissions vary considerably across latitudinal bands. In freshwater wetlands (except those between 20°S to 20°N) the spring onset of elevated CH4 emissions starts 3 d earlier, and the CH4 emission season lasts 4 d longer, for each degree Celsius increase in mean annual air temperature. On average, the spring onset of increasing CH4 emissions lags behind soil warming by 1 month, with very few sites experiencing increased CH4 emissions prior to the onset of soil warming. In contrast, roughly half of these sites experience the spring onset of rising CH4 emissions prior to the spring increase in gross primary productivity (GPP). The timing of peak summer CH4 emissions does not correlate with the timing for either peak summer temperature or peak GPP. Our results provide seasonality parameters for CH4 modeling and highlight seasonality metrics that cannot be predicted by temperature or GPP (i.e., seasonality of CH4 peak). FLUXNET-CH4 is a powerful new resource for diagnosing and understanding the role of terrestrial ecosystems and climate drivers in the global CH4 cycle, and future additions of sites in tropical ecosystems and site years of data collection will provide added value to this database. All seasonality parameters are available at https://doi.org/10.5281/zenodo.4672601 (Delwiche et al., 2021). Additionally, raw FLUXNET-CH4 data used to extract seasonality parameters can be downloaded from https://fluxnet.org/data/fluxnet-ch4-community-product/ (last access: 7 April 2021), and a complete list of the 79 individual site data DOIs is provided in Table 2 of this paper.

FLUXNET-CH4