Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Review and Analysis of Algorithmic Approaches Developed for Prognostics on CMAPSS Dataset

Benchmarking of prognostic algorithms has been challenging due to limited availability of common datasets suitable for prognostics. In an attempt to alleviate this problem several benchmarking datasets have been collected by NASA's prognostic center of excellence and made available to the Prognostics and Health Management (PHM) community to allow evaluation and comparison of prognostics algorithms. Among those datasets are five C-MAPSS datasets that have been extremely popular due to their unique characteristics making them suitable for prognostics. The C-MAPSS datasets pose several challenges that have been tackled by different methods in the PHM literature. In particular, management of high variability due to sensor noise, effects of operating conditions, and presence of multiple simultaneous fault modes are some factors that have great impact on the generalization capabilities of prognostics algorithms. More than 70 publications have used the C-MAPSS datasets for developing data-driven prognostic algorithms. The C-MAPSS datasets are also shown to be well-suited for development of new machine learning and pattern recognition tools for several key preprocessing steps such as feature extraction and selection, failure mode assessment, operating conditions assessment, health status estimation, uncertainty management, and prognostics performance evaluation. This paper summarizes a comprehensive literature review of publications using C-MAPSS datasets and provides guidelines and references to further usage of these datasets in a manner that allows clear and consistent comparison between different approaches.

Uncertainty↗

Total Ozone Trends from 1979 to 2016 Derived from Five Merged Observational Datasets - The Emergence into Ozone Recovery

We report on updated trends using different merged datasets from satellite and ground-based observations for the period from 1979 to 2016. Trends were determined by applying a multiple linear regression (MLR) to annual mean zonal mean data. Merged datasets used here include NASA MOD v8.6 and National Oceanic and Atmospheric Administration (NOAA) merge v8.6, both based on data from the series of Solar Backscatter UltraViolet (SBUV) and SBUV-2 satellite instruments (1978–present) as well as the Global Ozone Monitoring Experiment (GOME)-type Total Ozone (GTO) and GOME-SCIAMACHY-GOME-2 (GSG) merged datasets (1995-present), mainly comprising satellite data from GOME, the Scanning Imaging Absorption Spectrometer for Atmospheric Chartography (SCIAMACHY), and GOME-2A. The fifth dataset consists of the monthly mean zonal mean data from ground-based measurements collected at World Ozone and UV Data Center (WOUDC). The addition of four more years of data since the last World Meteorological Organization (WMO) ozone assessment (2013-2016) shows that for most datasets and regions the trends since the stratospheric halogen reached its maximum (approximately 1996 globally and approximately 2000 in polar regions) are mostly not significantly different from zero. However, for some latitudes, in particular the Southern Hemisphere extratropics and Northern Hemisphere subtropics, several datasets show small positive trends of slightly below +1 percent decade(exp. -1) that are barely statistically significant at the 2 Sigma uncertainty level. In the tropics, only two datasets show significant trends of +0.5 to +0.8 percent(exp.-1), while the others show near-zero trends. Positive trends since 2000 have been observed over Antarctica in September, but near-zero trends are found in October as well as in March over the Arctic. Uncertainties due to possible drifts between the datasets, from the merging procedure used to combine satellite datasets and related to the low sampling of ground-based data, are not accounted for in the trend analysis. Consequently, the retrieved trends can be only considered to be at the brink of becoming significant, but there are indications that we are about to emerge into the expected recovery phase. However, the recent trends are still considerably masked by the observed large year-to-year dynamical variability in total ozone.

column ozone trends↗

Land-Use Harmonization Datasets for Global Carbon Budget 2019 and Beyond

Land-use change has been the dominant source of anthropogenic carbon emissions for most of the historical period, and is currently one of the largest and most uncertain components of the global carbon cycle. Advancing the scientific understanding on this topic requires that the best data be used as input to the best models in well-organized scientific assessments. The Land-Use Harmonization dataset (LUH2), previously developed and used as input for CMIP6 simulations, has been updated annually to provide required input to land models in the annual Global Carbon Budget (GCB) assessment. These annual LUH2-GCB updates and extensions have incorporated annual FAO wood harvest data updates for dataset years after 2015 and HYDE gridded agriculture area data updates (based on annual FAO agricultural area data updates) for dataset years after 2012, along with extrapolations to the current year due to a lag of one or more years in the FAO data releases. The resulting updated LUH2-GCB datasets have provided global, annual gridded land-use and land-use change data relating toagricultural expansion, deforestation, wood harvesting, shifting cultivation, regrowth and afforestation, and crop rotationsand pasture managementand are used by both book-keeping models and Dynamic Global Vegetation Models (DGVMs) for the GCB. For GCB 2019,a more significant update to LUH2 was produced (LUH2-GCB2019) to correct cropland and grazing area errors in the underlying input datasets for the globally important region of Brazil, as far back as 1950. From 1951-2012 the LUH2-GCB2019 dataset begins to diverge from the LUH2 v2hdataset, with peak differences in Brazil in the year 2000 for grazing land (difference of 100,000 km2) and in the year 2009 for cropland (difference of 77,000 km2), along with significant sub-national reorganization of agricultural land-use patterns within Brazil. These LUH2-GCB2019 corrections for Brazil provide the base for future LUH2-GCB updates including the recent LUH2-GCB2020 dataset, and present a starting point for operationalizing the creation of these datasets toreduce time-lags due to the multiple input dataset and model latencies.

L Chini↗

Automated Collection of Scientific Publications Linked to NASA Earth Science Datasets

NASA's Earth Observing System Data and Information System (EOSDIS) began dataset Digital Object Identifier (DOI) registration in 2012. The number of dataset DOIs registered as of January of 2023 exceeds 11,000. As the research community becomes aware of the importance of sharing data through Open Science and optimizing data reuse through Findability, Accessibility, Interoperability, and Reuse (FAIR) data management principles, datasets are increasingly being cited in scientific publications. When datasets are cited explicitly by DOI within published works, automated methods can be developed for collecting these published works from a variety of bibliometric sources. The coverage of the sources varies, so each source can collect citations that are only available within it. Using major citation databases such as Scopus and Web of Science, the Google Scholar search engine, the CrossRef Open Citation Index, and the dataset DOI registry DataCite, we present an automated workflow for dataset citation collection. By harvesting citations automatically, a citation library is created explicitly linking EOSDIS datasets to publications that cite them. Using Zotero, a free and open-source citation manager, we demonstrate how to access and browse this library by the tags indicating bibliometric sources, dataset DOI, and the dataset archive center. We also demonstrate temporary trends in the number of publications harvested from bibliometric sources.

Infometrics↗

Creating A Consistent Historical NASA POWER Solar Radiation Dataset to Support Renewable Energy, Building Energy Efficiency and Agro-Climatology Decisions

Prediction of Worldwide Energy Resources (POWER) project provides irradiance dataset to support renewable energy, building energy efficiency and agricultural needs. These datasets are derived from Global Energy and Water Cycle Experiment Surface Radiation Budget (GEWEX SRB) and Clouds and the Earth’s Radiant Energy System (CERES SYN1Deg). A systematic bias has been reported between these two datasets for the years with overlapping observations. For obtaining a consistent climate data record spanning the entire time record of observations, it is crucial to understand and remove the bias in the irradiance dataset. Inconsistency in solar radiation data can lead to inaccurate conclusions about solar energy potential and obscure real trends in solar radiation patterns that would impact energy availability assessments. In this study, we adapt quantile mapping approach to remove the systematic bias and to improve reliability of shortwave and longwave irradiance data. We present a validation of the bias corrected data against ground truth. For each 1° latitude and 1° longitude grid box across the globe, we match the CDFs of the reference dataset (CERES SYN1Deg) to that of the SRB dataset, thereby, adjusting the irradiance values to match the empirical distribution of two different measurements. The performance of quantile mapping is evaluated by using the metrics such as Mean Absolute Deviation (MAD). The results indicate that the quantile mapping significantly improves the accuracy and reliability of solar irradiance dataset especially for the weather conditions associated with high cloud cover and extreme irradiance values. The initial range of MAD for the studied sites for daily data was 4 to 19 Wm-2. After correction these reduced to 3 to 7 Wm-2. The findings from this study have important implications for solar energy system design, agricultural planning, and climate modeling community. Reducing the inconsistency and biases in solar irradiance dataset can enable better planning and operation of solar energy systems, leading to increased efficiency and cost-effectiveness. Additionally, this work also contributes to the statistical post-processing techniques in the renewable energy domain and highlights the potential of historical and near-real-time NASA POWER dataset as a valuable resource for solar energy research and applications.

POWER↗

A comprehensive guide to CAN IDS data and introduction of the ROAD dataset

Although ubiquitous in modern vehicles, Controller Area Networks (CANs) lack basic security properties and are easily exploitable. A rapidly growing field of CAN security research has emerged that seeks to detect intrusions or anomalies on CANs. Producing vehicular CAN data with a variety of intrusions is a difficult task for most researchers as it requires expensive assets and deep expertise. To illuminate this task, we introduce the first comprehensive guide to the existing open CAN intrusion detection system (IDS) datasets. We categorize attacks on CANs including fabrication (adding frames, e.g., flooding or targeting and ID), suspension (removing an ID’s frames), and masquerade attacks (spoofed frames sent in lieu of suspended ones). We provide a quality analysis of each dataset; an enumeration of each datasets’ attacks, benefits, and drawbacks; categorization as real vs. simulated CAN data and real vs. simulated attacks; whether the data is raw CAN data or signal-translated; number of vehicles/CANs; quantity in terms of time; and finally a suggested use case of each dataset. State-of-the-art public CAN IDS datasets are limited to real fabrication (simple message injection) attacks and simulated attacks often in synthetic data, lacking fidelity. In general, the physical effects of attacks on the vehicle are not verified in the available datasets. Only one dataset provides signal-translated data but is missing a corresponding “raw” binary version. This issue pigeon-holes CAN IDS research into testing on limited and often inappropriate data (usually with attacks that are too easily detectable to truly test the method). The scarcity of appropriate data has stymied comparability and reproducibility of results for researchers. As our primary contribution, we present the Real ORNL Automotive Dynamometer (ROAD) CAN IDS dataset, consisting of over 3.5 hours of one vehicle’s CAN data. ROAD contains ambient data recorded during a diverse set of activities, and attacks of increasing stealth with multiple variants and instances of real (i.e. non-simulated) fuzzing, fabrication, unique advanced attacks, and simulated masquerade attacks. To facilitate a benchmark for CAN IDS methods that require signal-translated inputs, we also provide the signal time series format for many of the CAN captures. Our contributions aim to facilitate appropriate benchmarking and needed comparability in the CAN IDS research field.

97 MATHEMATICS AND COMPUTING↗

Dataset of U.S. School Bus Depots

A large body of public health literature describes how undesirable or dangerous facilities, such as truck depots and industrial plants, located in or near communities can lead to health harms. Research also describes the high levels of traffic-related air and noise pollution that is linked to health harms and may be disproportionately distributed near many schools. Therefore, a primary use case for this dataset is to analyze the location of school bus depots and to create an evidence base that would better enable the work of community members, advocates, and other stakeholders toward improving air quality and public health. Other possible uses for this school bus depot dataset include electricity grid planning and reliability, given recent momentum toward school bus electrification. This dataset was created using an object-based approach with remote sensing data. The primary source of aerial imagery was the National Agriculture Imagery Program (NAIP) dataset. NAIP imagery was analyzed to locate individual school buses based on their color and size, and then classified clusters of school buses as potential depots, which were then verified visually. The resulting dataset contains 11,309 depots across the 48 contiguous U.S. states and Washington, D.C. Fifty-one percent (5,730 depots) are at schools, defined as being 350 meters or less from the nearest school. The accuracy of the dataset was assessed by comparing it with independent reference datasets containing 506 depots from the records of two school transportation companies. We found good agreement, with an omission error rate of 15.2% (77 depots). This dataset represents one of the only remote sensing projects to conduct object detection using data at the sub-meter to 1-meter resolution for a continental-scale application.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ESS-DIVE Reporting Format for Dataset Package Metadata

ESS-DIVE’s (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) dataset metadata reporting format is intended to compile information about a dataset (e.g., title, description, funding sources) that can enable reuse of data submitted to the ESS-DIVE data repository. The files contained in this dataset include instructions (dataset_metadata_guide.md and README.md) that can be used to understand the types of metadata ESS-DIVE collects. The data dictionary (dd.csv) follows ESS-DIVE’s file-level metadata reporting format and includes brief descriptions about each element of the dataset metadata reporting format. This dataset also includes a terminology crosswalk (dataset_metadata_crosswalk.csv) that shows how ESS-DIVE’s metadata reporting format maps onto other existing metadata standards and reporting formats.Data contributors to ESS-DIVE can provide this metadata by manual entry using a web form or programmatically via ESS-DIVE’s API (Application Programming Interface). A metadata template (dataset_metadata_template.docx or dataset_metadata_template.pdf) can be used to collaboratively compile metadata before providing it to ESS-DIVE.Since being incorporated into ESS-DIVE’s data submission user interface, ESS-DIVE’s dataset metadata reporting format, has enabled features like automated metadata quality checks, and dissemination of ESS-DIVE datasets onto other data platforms including Google Dataset Search and DataCite.

54 ENVIRONMENTAL SCIENCES↗

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS↗

BUTTER-E - Energy Consumption Data for the BUTTER Empirical Deep Learning Dataset

The BUTTER-E - Energy Consumption Data for the BUTTER Empirical Deep Learning Dataset adds node-level energy consumption data from watt-meters to the primary sweep of the BUTTER - Empirical Deep Learning Dataset. This dataset contains energy consumption and performance data from 63,527 individual experimental runs spanning 30,582 distinct configurations: 13 datasets, 20 sizes (number of trainable parameters), 8 network "shapes", and 14 depths on both CPU and GPU hardware collected using node-level watt-meters. This dataset reveals the complex relationship between dataset size, network structure, and energy use, and highlights the impact of cache effects. BUTTER-E is intended to be joined with the BUTTER dataset (see "BUTTER - Empirical Deep Learning Dataset on OEDI" resource below) which characterizes the performance of 483k distinct fully connected neural networks but does not include energy measurements.

Array↗

A consistent dataset for the net income distribution for 190 countries and aggregated to 32 geographical regions from 1958 to 2015

Abstract. Data on income distributions within and across countries are becoming increasingly important for informing analysis of income inequality and understanding the distributional consequences of climate change. While datasets on income distribution collected from household surveys are available for multiple countries, these datasets often do not represent the same concept of inequality (or income concept) and therefore make comparisons across countries, over time and across datasets difficult. Here, we present a consistent dataset of income distributions across 190 countries from 1958 to 2015 measured in terms of net income. We complement the observed values in this dataset with values imputed from a summary measure of the income distribution, specifically the Gini coefficient. For the imputation, we use a recently developed nonparametric principal-component-based approach that shows an excellent fit to data on income distributions compared to other approaches. We also present another version of this dataset aggregated from the country level to 32 geographical regions. Our dataset is developed for the purpose of calibrating models such as integrated human–Earth system models with detailed data on income distributions. This dataset will enable more robust analysis of income distribution at multiple scales. The latest version of our data are available on Zenodo: https://doi.org/10.5281/zenodo.7093997 (Narayan et al., 2022b).

97 MATHEMATICS AND COMPUTING↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

A comparison between general circulation model simulations using two sea surface temperature datasets for January 1979

Simulations with the UCLA atmospheric general circulation model (AGCM) using two different global sea surface temperature (SST) datasets for January 1979 are compared. One of these datasets is based on Comprehensive Ocean-Atmosphere Data Set (COADS) (SSTs) at locations where there are ship reports, and climatology elsewhere; the other is derived from measurements by instruments onboard NOAA satellites. In the former dataset (COADS SST), data are concentrated along shipping routes in the Northern Hemisphere; in the latter dataset High Resolution Infrared Sounder (HIRS SST), data cover the global domain. Ensembles of five 30-day mean fields are obtained from integrations performed in the perpetual-January mode. The results are presented as anomalies, that is, departures of each ensemble mean from that produced in a control simulation with climatological SSTs. Large differences are found between the anomalies obtained using COADS and HIRS SSTs, even in the Northern Hemisphere where the datasets are most similar to each other. The internal variability of the circulation in the control simulation and the simulated atmospheric response to anomalous forcings appear to be linked in that the pattern of geopotential height anomalies obtained using COADS SSTs resembles the first empirical orthogonal function (EOF 1) in the control simulation. The corresponding pattern obtained using HIRS SSTs is substantially different and somewhat resembles EOF 2 in the sector from central North America to central Asia. To gain insight into the reasons for these results, three additional simulations are carried out with SST anomalies confined to regions where COADS SSTs are substantially warmer than HIRS SSTs. The regions correspond to warm pools in the northwest and northeast Pacific, and the northwest Atlantic. These warm pools tend to produce positive geopotential height anomalies in the northeastern part of the corresponding oceans. Both warm pools in the Pacific produce large-scale circulation anomalies with a pattern that resembles that obtained using COADS SSTs as well as EOF 1 of the control simulation; the warm pool in the Atlantic does not. These results suggest that the differences obtained with COADS SSTs and HIRS SSTs are mostly due to the differences in the datasets over the northern Pacific. There was a blocking episode near Greenland in late January 1979. Both simulations with warm SST anomalies over the northwest and northeast Pacific show a tendency toward increased incidence of North Atlantic blocking; the simulation with warm SST anomalies over the northwest Atlantic shows a tendency toward decreased incidence. These results suggest that features in both SST datasets that do not have a counterpart in the other dataset contribute signficantly to the differences between the simulated and observed fields. The results of this study imply that uncertainties in current SST distributions for the world oceans can be as important as the SST anomalies themselves in terms of their impact on the atmospheric circulation. Caution should be exercised, therefore, when linking anomalous circulation and SST patterns, especially in long-range prediction.

Ose, Tomoaki↗

The Transition of NASA EOS Datasets to WFO Operations: A Model for Future Technology Transfer

The collocation of a National Weather Service (NWS) Forecast Office with atmospheric scientists from NASA/Marshall Space Flight Center (MSFC) in Huntsville, Alabama has afforded a unique opportunity for science sharing and technology transfer. Specifically, the NWS office in Huntsville has interacted closely with research scientists within the SPORT (Short-term Prediction and Research and Transition) Center at MSFC. One significant technology transfer that has reaped dividends is the transition of unique NASA EOS polar orbiting datasets into NWS field operations. NWS forecasters primarily rely on the AWIPS (Advanced Weather Information and Processing System) decision support system for their day to day forecast and warning decision making. Unfortunately, the transition of data from operational polar orbiters or low inclination orbiting satellites into AWIPS has been relatively slow due to a variety of reasons. The ability to integrate these high resolution NASA datasets into operations has yielded several benefits. The MODIS (MODerate-resolution Imaging Spectrometer ) instrument flying on the Aqua and Terra satellites provides a broad spectrum of multispectral observations at resolutions as fine as 250m. Forecasters routinely utilize these datasets to locate fine lines, boundaries, smoke plumes, locations of fog or haze fields, and other mesoscale features. In addition, these important datasets have been transitioned to other WFOs for a variety of local uses. For instance, WFO Great Falls Montana utilizes the MODIS snow cover product for hydrologic planning purposes while several coastal offices utilize the output from the MODIS and AMSR-E instruments to supplement observations in the data sparse regions of the Gulf of Mexico and western Atlantic. In the short term, these datasets have benefited local WFOs in a variety of ways. In the longer term, the process by which these unique datasets were successfully transitioned to operations will benefit the planning and implementation of products and datasets derived from both NPP and NPOESS. This presentation will provide a brief overview of current WFO usage of satellite data, the transition of datasets between SPORT and the N W S , and lessons learned for future transition efforts.

Darden, C.↗

The SUMup Dataset: Compiled Measurements of Surface Mass Balance Components over Ice Sheets and Sea Ice with Analysis over Greenland

Increasing atmospheric temperatures over ice cover affect surface processes, including melt, snowfall, and snow density. Here, we present the Surface Mass Balance and Snow on Sea Ice Working Group (SUMup) dataset, a standardized dataset of Arctic and Antarctic observations of surface mass balance components. The July 2018 SUMup dataset consists of three subdatasets, snow/firn density (https://doi.org/10.18739/A2JH3D23R), at least near-annually resolved snow accumulation on land ice (https://doi.org/10.18739/A2DR2P790), and snow depth on sea ice (https://doi.org/10.18739/A2WS8HK6X), to monitor change and improve estimates of surface mass balance. The measurements in this dataset were compiled from field notes, papers, technical reports, and digital files. SUMup is a compiled, community-based dataset that can be and has been used to evaluate modeling efforts and remote sensing retrievals. Active submission of new or past measurements is encouraged. Analysis of the dataset shows that Greenland Ice Sheet density measurements in the top 1m do not show a strong relationship with annual temperature. At Summit Station, Greenland, accumulation and surface density measurements vary seasonally with lower values during summer months. The SUMup dataset is a dynamic, living dataset that will be updated and expanded for community use as new measurements are taken and new processes are discovered and quantified.

Montgomery, Lynn↗

Improving Earth Science Dataset Search with Publication

The NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) archives a large number of Earth observational datasets. Thousands of the publications are created each year based on these datasets. The content of these publications can be used for discovery of the datasets based on the characteristics of applicational research. We leverage the content of these publications to retrieve the information about phenomena and domains where measurements from the datasets were utilized through linking these publications and dataset in Knowledge Graph. We retrieve phenomena and domain information using SWEET ontology and produce the set of keywords that are linked to the datasets. Further, we evaluate this link strength according to the frequency of dataset usage in the papers mentioning these keywords. We demonstrate how this linkage can improve dataset search by comparing the search results obtained from Common Metadata Repository (CMR) search and the publications based data.

Kristina Stoyanova↗

Calculation of top-of-atmosphere, surface and atmospheric cloud radiative kernels and feedbacks based on ISCCP-H datasets

This study aims to create observation-based cloud radiative kernel (CRK) datasets and evaluate them by direct comparison of CRK and the CRK-derived cloud feedback datasets. Based on the International Satellite Cloud Climatology Project (ISCCP) H datasets, we calculate CRKs (called FH CRKs) as 2D joint function/histogram of cloud optical depth and cloud top pressure for shortwave, longwave, and their sum, Net, at the top of atmosphere (TOA), as well as, for the first time, at the surface (SFC) and in the atmosphere (ATM). The direct comparison shows that FH agrees reasonably well with three other TOA CRK datasets. With cloud fraction change (CFC) datasets of the same histogram for doubled-CO2 simulation from 10 CFMIP1 models, we derive all the TOA, SFC and ATM cloud feedback using the FH CRKs. Our TOA cloud feedback is highly similar to the previous counterparts. Based on the comparison for the 4 CRK datasets and the 10 CFC datasets, we estimate the uncertainty budget for the CRK-derived cloud feedback and show that the CFC-associated uncertainty contributes > 98.5% of the total cloud feedback uncertainty while CRK’s is very small. Our preliminary evaluation shows that some near-zero/small cloud feedback in the TOA-alone feedback indeed results from the compensation of sizable cloud feedback of the SFC and ATM feedback, demonstrating that they help reveal some significant surface and atmospheric cloud feedback whose sum appears insignificant in TOA-alone feedback

cloud radiative kernel (CRK) datasets↗

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source↗