Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

ResOpsUS, a dataset of historical reservoir operations in the contiguous United States

Abstract There are over 52,000 dams in the contiguous US ranging from 0.5 to 243 meters high that collectively hold 600,000 million cubic meters of water. These structures have dramatically affected the river dynamics of every major watershed in the country. While there are national datasets that document dam attributes, there is no national dataset of reservoir operations. Here we present a dataset of historical reservoir inflows, outflows and changes in storage for 679 major reservoirs across the US, called ResOpsUS. All of the data are provided at a daily temporal resolution. Temporal coverage varies by reservoir depending on construction date and digital data availability. Overall, the data spans from 1930 to 2020, although the best coverage is for the most recent years, particularly 1980 to 2020. The reservoirs included in our dataset cover more than half of the total storage of large reservoirs in the US (defined as reservoirs with storage greater 0.1 km 3 ). We document the assembly process of this dataset as well as its contents. Historical operations are also compared to static reservoir attribute datasets for validation.

54 ENVIRONMENTAL SCIENCES↗

Two excited-state datasets for quantum chemical UV-vis spectra of organic molecules

Abstract We present two open-source datasets that provide time-dependent density-functional tight-binding (TD-DFTB) electronic excitation spectra of organic molecules. These datasets represent predictions of UV-vis absorption spectra performed on optimized geometries of the molecules in their electronic ground state. The GDB-9-Ex dataset contains a subset of 96,766 organic molecules from the original open-source GDB-9 dataset. The ORNL_AISD-Ex dataset consists of 10,502,904 organic molecules that contain between 5 and 71 non-hydrogen atoms. The data reveals the close correlation between the magnitude of the gaps between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO), and the excitation energy of the lowest singlet excited state energies quantitatively. The chemical variability of the large number of molecules was examined with a topological fingerprint estimation based on extended-connectivity fingerprints (ECFPs) followed by uniform manifold approximation and projection (UMAP) for dimension reduction. Both datasets were generated using the DFTB+ software on the “Andes” cluster of the Oak Ridge Leadership Computing Facility (OLCF).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Typical and extreme weather datasets for studying the resilience of buildings to climate change and heatwaves

We present unprecedented datasets of current and future projected weather files for building simulations in 15 major cities distributed across 10 climate zones worldwide. The datasets include ambient air temperature, relative humidity, atmospheric pressure, direct and diffuse solar irradiance, and wind speed at hourly resolution, which are essential climate elements needed to undertake building simulations. The datasets contain typical and extreme weather years in the EnergyPlus weather file (EPW) format and multiyear projections in comma-separated value (CSV) format for three periods: historical (2001–2020), future mid-term (2041–2060), and future long-term (2081–2100). The datasets were generated from projections of one regional climate model, which were bias-corrected using multiyear observational data for each city. The methodology used makes the datasets among the first to incorporate complex changes in the future climate for the frequency, duration, and magnitude of extreme temperatures. These datasets, created within the IEA EBC Annex 80 “Resilient Cooling for Buildings”, are ready to be used for different types of building adaptation and resilience studies to climate change and heatwaves.

54 ENVIRONMENTAL SCIENCES↗

Search for ultralight dark matter in the SuperMAG high-fidelity dataset

Ultralight dark matter, such as kinetically mixed dark-photon dark matter (DPDM) or axion-like-particle dark matter (axion DM), can source an oscillating magnetic-field signal at Earth’s surface. Previous work searched for this signal in a publicly available dataset of global magnetometer measurements maintained by the SuperMAG collaboration. This “low-fidelity” dataset reported measurements with a 1-min time resolution, allowing the search to set leading direct constraints on DPDM and axion DM with Compton frequencies f DM ≤ 1 / ( 1 min ) (corresponding to masses m DM ≤ 7 × 10 − 17 eV ). More recently, a dedicated experiment undertaken by the SNIPE Hunt collaboration has also searched for this same signal at higher frequencies f DM ≥ 0.5 Hz (or m DM ≥ 2 × 10 − 15 eV ). In this work, we search for this signal of ultralight DM in the SuperMAG “high-fidelity” dataset, which features a 1-sec time resolution, allowing us to probe the gap in parameter space between the low-fidelity dataset and the SNIPE Hunt experiment. The high-fidelity dataset exhibits lower geomagnetic noise than the low-fidelity dataset and features more data than the SNIPE Hunt experiment, making it a powerful probe of ultralight DM. Our search finds no robust DPDM or axion DM candidates. We set constraints on DPDM and axion DM parameter space for 10 − 3 Hz ≤ f DM ≤ 0.98 Hz (or 4 × 10 − 18 eV ≤ m DM ≤ 4 × 10 − 15 eV ). Our results are the leading direct constraints on both DPDM and axion DM in this mass range, and our DPDM constraint surpasses the leading astrophysical constraint in a narrow range around m A ′ ≈ 2 × 10 − 15 eV . Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data Fusion for the Development of a Multimodal Freight Transload Facilities Dataset in the U.S.

To withstand the growing demand of commodity volume and its strain on the transportation infrastructure, it is necessary to identify the flow of commodities by route and mode. However, a national multimodal freight routing model does not exist for the U.S. The development of such model requires multiple building blocks, such as virtual representations of roadway, railway, and waterway networks, transload facilities (TFs), and access/egress links. Most of these blocks have a robust database in the U.S., except for the TFs. Here, this paper presents the fusion of dispersed and heterogeneous representations of multimodal TFs into a single, comprehensive, geospatial freight TF dataset. The TF dataset is derived from several sources, including the U.S. Army Corps of Engineers Master Docks Plus, the National Transportation Atlas Database, the Intermodal Association of North America, industry publications, and other public information. First, individual datasets were queried and reconciled. A geocoding/reverse geocoding process was applied to get the best street address and latitude/longitude location for each terminal. Then, duplicate terminals were identified by a fuzzy match algorithm based on terminal name and location, and removed. Validation was performed by visual inspection of random facilities. The main contributions of this work are: a publicly available version of the TF dataset, including facility location and multimodal transfer capability of 9,003 facilities, and an enterprise-version with the same facilities but including commodity handling capabilities. The main purpose of developing the TF dataset is to inform multimodal routing algorithms. The proposed TF dataset allows for credibly modeling the multimodal transfer of commodities within shipment routes.

Commodity Routing↗

2021 Smoky Mountains Conference Data Challenge Synthetic-to-Real Domain Adaptation for Autonomous Driving Dataset

The dataset is comprised of both real and synthetic images from a vehicle's forward-facing camera. Each camera image is accompanied by a corresponding pixel-level semantic segmentation image (all files are .png files). In total, the dataset contains 5600 images in the training/validation set and 1400 images in the testing set. The training dataset contains mostly synthetic RGB images collected with a wide range of weather and lighting conditions using the CARLA simulator [1]. In addition, the training data also includes a small pre-selected subset of data from the Cityscapes training dataset – which is comprised of RGB-segmentation image pairs from driving scenarios in various European cities [2]. The testing data is split into three sets. The first set contains synthetic CARLA images with weather/lighting conditions that were not present in the training set. The second set is a subset of the Cityscapes testing dataset. Finally, the third set is an unknown testing set which will not be revealed to the participants until after the submission deadline. [1] Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, October). CARLA: An open urban driving simulator. In Conference on robot learning (pp. 1-16). PMLR. [2] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... and Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213-3223).

99 GENERAL AND MISCELLANEOUS↗

DEEPEN 3D PFA Weights for Exploration Datasets in Magmatic Environments

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This GDR submission includes those weights. The weighting was done using two different approaches: one based on expert opinions, and one based on statistical learning. The weights are intended to describe how useful a particular exploration method is for imaging each component of each play type. They may be adjusted based on the characteristics of the resource under investigation, knowledge of the quality of the dataset, or simply to reduce the impact a single dataset has on the resulting outputs. Within the DEEPEN PFA, separate sets of weights are produced for each component of each play type, since exploration methods hold different levels of importance for detecting each play component, within each play type. The weights for conventional hydrothermal systems were based on the average of the normalized weights used in the DOE-funded PFA projects that were focused on magmatic plays. This decision was made because conventional hydrothermal plays are already well-studied and understood, and therefore it is logical to use existing weights where possible. In contrast, a true PFA has never been applied to superhot EGS or supercritical plays, meaning that exploration methods have never been weighted in terms of their utility in imaging the components of these plays. To produce weights for superhot EGS and supercritical plays, two different approaches were used: one based on expert opinion and the analytical hierarchy process (AHP), and another using a statistical approach based on principal component analysis (PCA). The weights are intended to provide standardized sets of weights for each play type in all magmatic geothermal systems. Two different approaches were used to investigate whether a more data-centric approach might allow new insights into the datasets, and also to analyze how different weighting approaches impact the outcomes. The expert/AHP approach involved using an online tool (https://bpmsg.com/ahp/) with built-in forms to make pairwise comparisons which are used to rank exploration methods against one-another. The inputs are then combined in a quantitative way, ultimately producing a set of consensus-based weights. To minimize the burden on each individual participant, the forms were completed in group discussions. While the group setting means that there is potential for some opinions to outweigh others, it also provides a venue for conversation to take place, in theory leading the group to a more robust consensus then what can be achieved on an individual basis. This exercise was done with two separate groups: one consisting of U.S.-based experts, and one consisting of Iceland-based experts in magmatic geothermal systems. The two sets of weights were then averaged to produce what we will from here on refer to as the "expert opinion-based weights," or "expert weights" for short. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. More information on this approach along with the dataset used to produce the statistical weights may be found in the linked dataset below.

15 GEOTHERMAL ENERGY↗

QA/QC of the East River, Colorado, discharge and geochemical time series datasets (Almont, BCC, and Pump House) to be used for modeling of hydrogeochemical balance

The following datasets were QA/QC-ed (Quality Assurance/Quality Control): 1. Brush Creek Confluence (BCC) discharge data (from Helen Malenda, USGS, Colorado School of Mines), which were calculated using the pressure transducer data and rating curves. The original 15 min time series data were presented as mean daily discharge. 2. Almont discharge data from United States Geological Survey (USGS). The original data were in 15 min time intervals, and were averaged to mean daily discharge time series. 3. Pump House discharge data as mean daily discharge (downloaded from the SFA portal). 4. BCC and Pump House chemistry data from SFA data portal and/or original spreadsheets provided by Roelof Versteeg. The following challenging QA/QC problems of the datasets were resolved: Missing data with the duration of gaps up to >1 month; Duplicated dates; Anomalies and outliers of discharge and concentrations; Time stamps of measurements of the discharge and concentrations are not aligned (hydrogeochemical balance calculations require the timestamps to be aligned). All QA/QC-ed datasets are given as csv files. The csv files were prepared using the xts files with multiple worksheets, which are also included in the data packages. Figures of the QA/QC-ed datasets are given in the jpeg and pdf formats. The QA/QC-ed datasets have been used to quantify discharge and chemical concentrations in river water in order to understand riverine exports of water and dissolved constituents in the East River watershed. These datasets served as a basis in the presentation given by P. Fox et al. at the 2021 Goldschmidt Conference.

54 ENVIRONMENTAL SCIENCES↗

U.S. Freight Transload Facilities Dataset

The U.S. Freight Transload Facilities Dataset provides location information (latitude, longitude, zip, city, county, state)for more than 9,000 facilities across 50 U.S. States where freight may be transferred between waterways, railways, and roadways. The dataset lists the known modes and available direction(s) for freight transfers at each facility as of 2024. The U.S. Freight Transload Facilities dataset was built by mining and fusing several public sources, such as the USACE Master Docks Plus, the USDOT National Transportation Atlas Database (NTAD), files from the Intermodal Association of North America (IANA), and the industry publication Bulk Transloader. The dataset constitutes a key piece of a multimodal freight transportation network and routing algorithm developed by USACE-ERDC. The dataset is shared as a .csv file. The dataset is published for research purposes and should not be considered exhaustive or authoritative.

Peterson, Steven [ORNL] (ORCID:0000000287672998)↗

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

The executive disruption model of tinnitus distress: Model validation in two independent datasets using factor score regression

This study presents the executive disruption model (EDM) of tinnitus distress and subsequently validates it statistically using two independent datasets (the Construction Dataset: n = 96 and the Validation Dataset: n = 200). The conceptual EDM was first operationalised as a structural causal model (construction phase). Then multiple regression was used to examine the effect of executive functioning on tinnitus-related distress (validation phase), adjusting for the additional contributions of hearing threshold and psychological distress. For both datasets, executive functioning negatively predicted tinnitus distress score by a similar amount (the Construction Dataset: β = −3.50, p = 0.13 and the Validation Dataset: β = −3.71, p = 0.02). Theoretical implications and applications of the EDM are subsequently discussed; these include the predictive nature of executive functioning in the development of distressing tinnitus, and the clinical utility of the EDM.

Clarke, Nathan A.↗

Panta Rhei benchmark dataset: socio-hydrological data of paired events of floods and droughts

As the adverse impacts of hydrological extremes increase in many regions of the world, a better understanding of the drivers of changes in risk and impacts is essential for effective flood and drought risk management and climate adaptation. However, there is currently a lack of comprehensive, empirical data about the processes, interactions, and feedbacks in complex human–water systems leading to flood and drought impacts. Here we present a benchmark dataset containing socio-hydrological data of paired events, i.e. two floods or two droughts that occurred in the same area. The 45 paired events occurred in 42 different study areas and cover a wide range of socio-economic and hydro-climatic conditions. The dataset is unique in covering both floods and droughts, in the number of cases assessed and in the quantity of socio-hydrological data. The benchmark dataset comprises (1) detailed review-style reports about the events and key processes between the two events of a pair; (2) the key data table containing variables that assess the indicators which characterize management shortcomings, hazard, exposure, vulnerability, and impacts of all events; and (3) a table of the indicators of change that indicate the differences between the first and second event of a pair. The advantages of the dataset are that it enables comparative analyses across all the paired events based on the indicators of change and allows for detailed context- and location-specific assessments based on the extensive data and reports of the individual study areas. The dataset can be used by the scientific community for exploratory data analyses, e.g. focused on causal links between risk management; changes in hazard, exposure and vulnerability; and flood or drought impacts. The data can also be used for the development, calibration, and validation of socio-hydrological models. The dataset is available to the public through the GFZ Data Services (Kreibich et al., 2023, https://doi.org/10.5880/GFZ.4.4.2023.001).

54 ENVIRONMENTAL SCIENCES↗

ClimateNet: an expert-labeled open dataset and deep learning architecture for enabling high-precision analyses of extreme weather

Abstract. Identifying, detecting, and localizing extreme weather events is a crucial first step in understanding how they may vary under different climate change scenarios. Pattern recognition tasks such as classification, object detection, and segmentation (i.e., pixel-level classification) have remained challenging problems in the weather and climate sciences. While there exist many empirical heuristics for detecting extreme events, the disparities between the output of these different methods even for a single event are large and often difficult to reconcile. Given the success of deep learning (DL) in tackling similar problems in computer vision, we advocate a DL-based approach. DL, however, works best in the context of supervised learning – when labeled datasets are readily available. Reliable labeled training data for extreme weather and climate events is scarce. We create “ClimateNet” – an open, community-sourced human-expert-labeled curated dataset that captures tropical cyclones (TCs) and atmospheric rivers (ARs) in high-resolution climate model output from a simulation of a recent historical period. We use the curated ClimateNet dataset to train a state-of-the-art DL model for pixel-level identification – i.e., segmentation – of TCs and ARs. We then apply the trained DL model to historical and climate change scenarios simulated by the Community Atmospheric Model (CAM5.1) and show that the DL model accurately segments the data into TCs, ARs, or “the background” at a pixel level. Further, we show how the segmentation results can be used to conduct spatially and temporally precise analytics by quantifying distributions of extreme precipitation conditioned on event types (TC or AR) at regional scales. The key contribution of this work is that it paves the way for DL-based automated, high-fidelity, and highly precise analytics of climate data using a curated expert-labeled dataset – ClimateNet. ClimateNet and the DL-based segmentation method provide several unique capabilities: (i) they can be used to calculate a variety of TC and AR statistics at a fine-grained level; (ii) they can be applied to different climate scenarios and different datasets without tuning as they do not rely on threshold conditions; and (iii) the proposed DL method is suitable for rapidly analyzing large amounts of climate model output. While our study has been conducted for two important extreme weather patterns (TCs and ARs) in simulation datasets, we believe that this methodology can be applied to a much broader class of patterns and applied to observational and reanalysis data products via transfer learning.

54 ENVIRONMENTAL SCIENCES↗

Benchmark datasets for SARS-CoV-2 surveillance bioinformatics

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the cause of coronavirus disease 2019 (COVID-19), has spread globally and is being surveilled with an international genome sequencing effort. Surveillance consists of sample acquisition, library preparation, and whole genome sequencing. This has necessitated a classification scheme detailing Variants of Concern (VOC) and Variants of Interest (VOI), and the rapid expansion of bioinformatics tools for sequence analysis. These bioinformatic tools are means for major actionable results: maintaining quality assurance and checks, defining population structure, performing genomic epidemiology, and inferring lineage to allow reliable and actionable identification and classification. Additionally, the pandemic has required public health laboratories to reach high throughput proficiency in sequencing library preparation and downstream data analysis rapidly. However, both processes can be limited by a lack of a standardized sequence dataset. We identified six SARS-CoV-2 sequence datasets from recent publications, public databases and internal resources. In addition, we created a method to mine public databases to identify representative genomes for these datasets. Using this novel method, we identified several genomes as either VOI/VOC representatives or non-VOI/VOC representatives. To describe each dataset, we utilized a previously published datasets format, which describes accession information and whole dataset information. Additionally, a script from the same publication has been enhanced to download and verify all data from this study.

60 APPLIED LIFE SCIENCES↗

Heuristics for Relevancy Ranking of Earth Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗

Relevancy Ranking of Satellite Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗

Development of a Knowledge Graph for Dataset Discovery and Identification at a NASA Data Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) archives and distributes hundreds of Earth Science data collections to the public. These collections are used in research, resulting in the publication of thousands of scientific papers each year. As new users come to GES DISC for data, it is important for them to understand how prior research used the data. To help researchers, a knowledge graph (KG) was designed and implemented to connect publication citations with dataset metadata. The relationships created in the graph have the potential to allow the Web applications that utilize this information to directly connect the publication to the GES DISC datasets and services. These relationships are demonstrated using a web application prototype. In addition, the graph can also make connections between publications, datasets, and measurements based on the mentions of datasets and their attributes in the publications. To demonstrate this capability, a web application was created that takes the excerpt from the publication and returns a most likely dataset and measurement pairing, ranking the results based on how often these datasets and measurements were used in prior publications.

Nathaniel Crosby↗