Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A globally sampled high-resolution hand-labeled validation dataset for evaluating surface water extent maps

Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.

54 ENVIRONMENTAL SCIENCES↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

Meta-analysis of North American Arctic and boreal aboveground biomass datasets: assessing accuracy, dynamics, and similarities

The North American arctic and boreal regions (ABRs) are rapidly warming and experiencing intensifying disturbances. Accurately quantifying aboveground biomass (AGB) is critical for understanding the impacts of these changes on the carbon cycle and for designing climate change mitigation strategies. Several AGB maps have been developed for the North American ABRs, including recent contributions from National Aeronautics and Space Administration’s Arctic-Boreal Vulnerability Experiment (ABoVE) campaign. However, these maps differ widely in training data, methodology, and resulting AGB density estimates. Presently, a comprehensive comparative evaluation is lacking, making it difficult for users to select datasets suited to their research or management needs. Here, in this study, we conducted a comparative analysis of nine AGB density datasets across North American ABRs, specifically for Alaska and Canada. We (1) summarized AGB by ecoregion and Canadian provinces, (2) evaluated their accuracy against field-based measurements, (3) analyzed spatial and temporal similarities among datasets, and (4) assessed their ability to capture disturbance (fire and harvest) impacts on AGB. We found substantial variation in regional and local AGB estimates across datasets, with overall accuracy ranging from R 2 = 0.25–0.62 and Bias% from −47.8% to 69.9% when validated against field plots. Despite these differences, most datasets have comparatively consistent spatial patterns in AGB (r > 0.8 for most cases). In contrast, agreement on the temporal patterns of AGB change is generally low. We found datasets with spatial resolutions ⩽300 m are capable of capturing disturbance impacts on AGB dynamics, though sensitivity varies across products. Our findings and dataset summary provide guidance for selecting appropriate AGB datasets for different applications within our study area. Our analysis also highlights the need to decrease map bias and increase capability to detect temporal change to decrease uncertainty of AGB datasets potentially by using training data which is representative of major plant functional types within the mapped area.

ABoVE↗

Land-use harmonization datasets for annual global carbon budgets

Abstract. Land-use change has been the dominant source of anthropogenic carbon emissions for most of the historical period and is currently one of the largest and most uncertain components of the global carbon cycle. Advancing the scientific understanding on this topic requires that the best data be used as input to state-of-the-art models in well-organized scientific assessments. The Land-Use Harmonization 2 dataset (LUH2), previously developed and used as input for simulations of the 6th Coupled Model Intercomparison Project (CMIP6), has been updated annually to provide required input to land models in the annual Global Carbon Budget (GCB) assessments. Here we discuss the methodology for producing these annual LUH2-GCB updates and extensions which incorporate annual wood harvest data updates from the Food and Agriculture Organization (FAO) of the United Nations for dataset years after 2015 and the History Database of the Global Environment (HYDE) gridded cropland and grazing area data updates (based on annual FAO cropland and grazing area data updates) for dataset years after 2012, along with extrapolations to the current year due to a lag of 1 or more years in the FAO data releases. The resulting updated LUH2-GCB datasets have provided global, annual gridded land-use and land-use-change data relating to agricultural expansion, deforestation, wood harvesting, shifting cultivation, regrowth and afforestation, crop rotations, and pasture management and are used by both bookkeeping models and dynamic global vegetation models (DGVMs) for the GCB. For GCB 2019, a more significant update to LUH2 was produced, LUH2-GCB2019 (https://doi.org/10.3334/ORNLDAAC/1851, Chini et al., 2020b), to take advantage of new data inputs that corrected cropland and grazing areas in the globally important region of Brazil as far back as 1950. From 1951 to 2012 the LUH2-GCB2019 dataset begins to diverge from the version of LUH2 used for the World Climate Research Programme's CMIP6, with peak differences in Brazil in the year 2000 for grazing land (difference of 100 000 km2) and in the year 2009 for cropland (difference of 77 000 km2), along with significant sub-national reorganization of agricultural land-use patterns within Brazil. The LUH2-GCB2019 dataset provides the base for future LUH2-GCB updates, including the recent LUH2-GCB2020 dataset, and presents a starting point for operationalizing the creation of these datasets to reduce time lags due to the multiple input dataset and model latencies.

Chini, Louise (ORCID:0000000290703505)↗

Observational ozone datasets over the global oceans and polar regions (version 2024)

Studying tropospheric ozone over the remote areas of the planet, such as the open oceans and the polar regions, is crucial to understand the role of ozone as a global climate forcer and regulator of atmospheric oxidative capacity. A focus on the pristine oceanic and polar regions complements the available land-based datasets and provides insights into key photochemical and depositional loss processes that control the concentrations and spatiotemporal variability in ozone as well as the physicochemical mechanisms driving these patterns. However, an assessment of the role of ozone over the oceanic and polar regions has been hampered by a lack of comprehensive observational datasets. Here, we present the first comprehensive collection of ozone data over the oceans and the polar regions. The overall dataset consists of 77 ship cruises/buoy-based observations and 48 aircraft-based campaigns. The dataset, consisting of more than 630 000 independent ozone measurement data points covering the period from 1977 to 2022 and an altitude range from the surface to 5000 m (with a focus on the lowest 2000 m), allows systematic analyses of the spatiotemporal distribution and long-term trends over the 11 defined ocean/polar regions. The datasets from ships, buoys, and aircraft are complemented by ozonesonde data from 29 launch sites or field campaigns and by 21 non-polar and 17 polar ground-based station datasets. The datasets contain information on how long the observed air masses were isolated from land, as estimated by backward trajectories from the individual observation points. To extract observations representative of oceanic conditions, we recommend using a subset of the data with an isolation time of 72 h or longer, from the analysis with coincident radon observations. These filtered oceanic and polar data showed typically flat diurnal cycles at high latitudes, whereas daytime decreases in ozone (11 %–16 %) were observed at lower latitudes. The ship/buoy- and aircraft-based datasets presented here will supplement the land-based ones in the TOAR-II (Tropospheric Ozone Assessment Report Phase II) database to provide a fully global assessment of tropospheric ozone. The described dataset is available at https://doi.org/10.17596/0004044 (Kanaya et al., 2025).

Kanaya, Yugo [Japan Agency for Marine-Earth Scienc↗

Performance of wind assessment datasets in United States coastal areas

The atmospheric dynamics that occur near the intersection of land and water offer exciting and challenging opportunities for wind energy deployment in coastal locations. New models and tools are continually being developed in support of wind resource assessment, and three recent products are explored in this work for their performance in representing characteristics of the wind resource at coastal locations: the Global Wind Atlas 3 (GWA3), the 2023 National Offshore Wind dataset (NOW-23), and the wind climate simulations that are a component of the Wind Integration National Dataset (WIND) Toolkit Long-Term Ensemble Dataset (WTK-LED Climate). These relatively new products are freely available and user-friendly so that anyone – from a utility-scale developer to a resident or business owner – can evaluate the potential for wind energy generation at their location of interest. The validations in this work provide guidance on the accuracy of wind resource assessments for coastal customers interested in installing small or midsize wind turbines (≤ 1 MW in capacity) to support energy needs at the residential, business, or community scale, such as the island and remotely located participants of the U.S. Department of Energy's Energy Transitions Initiative Partnership Project. At 23 coastal locations across the United States, dataset performance varies according to different evaluation metrics. All three recent datasets tend to overestimate the observed coastal wind resource. GWA3 produces the smallest annual average wind speed relative errors, whereas WTK-LED Climate is in best agreement in terms of representing diurnal wind speed cycles. NOW-23 is the highest performing of the datasets for representing seasonal and interannual trends in the coastal wind resource. While GWA3 and WTK-LED Climate are relatively insensitive to the dataset output heights selected for wind resource assessment at small and midsize wind turbine hub heights (20–60 m), significant variation in the NOW-23 representation of wind shear across the wind profile in the lowest 100 m of the atmosphere leads to notable differences in wind speed estimates according to the dataset output heights selected for evaluation. GWA3 exhibits challenges in the representation of observed wind speed diurnal cycles at small and midsize turbine hub heights, likely due to the dataset's consistent treatment of hourly wind speed trends regardless of altitude.

17 WIND ENERGY↗

OPFLearnData: Dataset for Learning AC Optimal Power Flow

The datasets are resulting from OPFLearn.jl, a Julia package for creating AC OPF datasets. The package was developed to provide researchers with a standardized way to efficiently create AC OPF datasets that are representative of more of the AC OPF feasible load space compared to typical dataset creation methods. The OPFLearn dataset creation method uses a relaxed AC OPF formulation to reduce the volume of the unclassified input space throughout the dataset creation process. The dataset contains load profiles and their respective optimal primal and dual solutions. Load samples are processed using AC OPF formulations from PowerModels.jl. More information on the dataset creation method can be found in our publication, "OPF-Learn: An Open-Source Framework for Creating Representative AC Optimal Power Flow Datasets" and in the package website: https://github.com/NREL/OPFLearn.jl.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A comprehensive guide to CAN IDS data and introduction of the ROAD dataset

Although ubiquitous in modern vehicles, Controller Area Networks (CANs) lack basic security properties and are easily exploitable. A rapidly growing field of CAN security research has emerged that seeks to detect intrusions or anomalies on CANs. Producing vehicular CAN data with a variety of intrusions is a difficult task for most researchers as it requires expensive assets and deep expertise. To illuminate this task, we introduce the first comprehensive guide to the existing open CAN intrusion detection system (IDS) datasets. We categorize attacks on CANs including fabrication (adding frames, e.g., flooding or targeting and ID), suspension (removing an ID’s frames), and masquerade attacks (spoofed frames sent in lieu of suspended ones). We provide a quality analysis of each dataset; an enumeration of each datasets’ attacks, benefits, and drawbacks; categorization as real vs. simulated CAN data and real vs. simulated attacks; whether the data is raw CAN data or signal-translated; number of vehicles/CANs; quantity in terms of time; and finally a suggested use case of each dataset. State-of-the-art public CAN IDS datasets are limited to real fabrication (simple message injection) attacks and simulated attacks often in synthetic data, lacking fidelity. In general, the physical effects of attacks on the vehicle are not verified in the available datasets. Only one dataset provides signal-translated data but is missing a corresponding “raw” binary version. This issue pigeon-holes CAN IDS research into testing on limited and often inappropriate data (usually with attacks that are too easily detectable to truly test the method). The scarcity of appropriate data has stymied comparability and reproducibility of results for researchers. As our primary contribution, we present the Real ORNL Automotive Dynamometer (ROAD) CAN IDS dataset, consisting of over 3.5 hours of one vehicle’s CAN data. ROAD contains ambient data recorded during a diverse set of activities, and attacks of increasing stealth with multiple variants and instances of real (i.e. non-simulated) fuzzing, fabrication, unique advanced attacks, and simulated masquerade attacks. To facilitate a benchmark for CAN IDS methods that require signal-translated inputs, we also provide the signal time series format for many of the CAN captures. Our contributions aim to facilitate appropriate benchmarking and needed comparability in the CAN IDS research field.

97 MATHEMATICS AND COMPUTING↗

Dataset of U.S. School Bus Depots

A large body of public health literature describes how undesirable or dangerous facilities, such as truck depots and industrial plants, located in or near communities can lead to health harms. Research also describes the high levels of traffic-related air and noise pollution that is linked to health harms and may be disproportionately distributed near many schools. Therefore, a primary use case for this dataset is to analyze the location of school bus depots and to create an evidence base that would better enable the work of community members, advocates, and other stakeholders toward improving air quality and public health. Other possible uses for this school bus depot dataset include electricity grid planning and reliability, given recent momentum toward school bus electrification. This dataset was created using an object-based approach with remote sensing data. The primary source of aerial imagery was the National Agriculture Imagery Program (NAIP) dataset. NAIP imagery was analyzed to locate individual school buses based on their color and size, and then classified clusters of school buses as potential depots, which were then verified visually. The resulting dataset contains 11,309 depots across the 48 contiguous U.S. states and Washington, D.C. Fifty-one percent (5,730 depots) are at schools, defined as being 350 meters or less from the nearest school. The accuracy of the dataset was assessed by comparing it with independent reference datasets containing 506 depots from the records of two school transportation companies. We found good agreement, with an omission error rate of 15.2% (77 depots). This dataset represents one of the only remote sensing projects to conduct object detection using data at the sub-meter to 1-meter resolution for a continental-scale application.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ESS-DIVE Reporting Format for Dataset Package Metadata

ESS-DIVE’s (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) dataset metadata reporting format is intended to compile information about a dataset (e.g., title, description, funding sources) that can enable reuse of data submitted to the ESS-DIVE data repository. The files contained in this dataset include instructions (dataset_metadata_guide.md and README.md) that can be used to understand the types of metadata ESS-DIVE collects. The data dictionary (dd.csv) follows ESS-DIVE’s file-level metadata reporting format and includes brief descriptions about each element of the dataset metadata reporting format. This dataset also includes a terminology crosswalk (dataset_metadata_crosswalk.csv) that shows how ESS-DIVE’s metadata reporting format maps onto other existing metadata standards and reporting formats.Data contributors to ESS-DIVE can provide this metadata by manual entry using a web form or programmatically via ESS-DIVE’s API (Application Programming Interface). A metadata template (dataset_metadata_template.docx or dataset_metadata_template.pdf) can be used to collaboratively compile metadata before providing it to ESS-DIVE.Since being incorporated into ESS-DIVE’s data submission user interface, ESS-DIVE’s dataset metadata reporting format, has enabled features like automated metadata quality checks, and dissemination of ESS-DIVE datasets onto other data platforms including Google Dataset Search and DataCite.

54 ENVIRONMENTAL SCIENCES↗

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS↗

BUTTER-E - Energy Consumption Data for the BUTTER Empirical Deep Learning Dataset

The BUTTER-E - Energy Consumption Data for the BUTTER Empirical Deep Learning Dataset adds node-level energy consumption data from watt-meters to the primary sweep of the BUTTER - Empirical Deep Learning Dataset. This dataset contains energy consumption and performance data from 63,527 individual experimental runs spanning 30,582 distinct configurations: 13 datasets, 20 sizes (number of trainable parameters), 8 network "shapes", and 14 depths on both CPU and GPU hardware collected using node-level watt-meters. This dataset reveals the complex relationship between dataset size, network structure, and energy use, and highlights the impact of cache effects. BUTTER-E is intended to be joined with the BUTTER dataset (see "BUTTER - Empirical Deep Learning Dataset on OEDI" resource below) which characterizes the performance of 483k distinct fully connected neural networks but does not include energy measurements.

Array↗

A consistent dataset for the net income distribution for 190 countries and aggregated to 32 geographical regions from 1958 to 2015

Abstract. Data on income distributions within and across countries are becoming increasingly important for informing analysis of income inequality and understanding the distributional consequences of climate change. While datasets on income distribution collected from household surveys are available for multiple countries, these datasets often do not represent the same concept of inequality (or income concept) and therefore make comparisons across countries, over time and across datasets difficult. Here, we present a consistent dataset of income distributions across 190 countries from 1958 to 2015 measured in terms of net income. We complement the observed values in this dataset with values imputed from a summary measure of the income distribution, specifically the Gini coefficient. For the imputation, we use a recently developed nonparametric principal-component-based approach that shows an excellent fit to data on income distributions compared to other approaches. We also present another version of this dataset aggregated from the country level to 32 geographical regions. Our dataset is developed for the purpose of calibrating models such as integrated human–Earth system models with detailed data on income distributions. This dataset will enable more robust analysis of income distribution at multiple scales. The latest version of our data are available on Zenodo: https://doi.org/10.5281/zenodo.7093997 (Narayan et al., 2022b).

97 MATHEMATICS AND COMPUTING↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DEEPEN Leapfrog Geodata Model Cleaned and Reformatted Exploration Datasets from Newberry Volcano

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the DEEPEN 3D play fairway analysis (PFA) conducted at Newberry Volcano for multiple play types (conventional hydrothermal, superhot EGS, and supercritical), existing geoscientific exploration datasets needed to be acquired, cleaned, reformatted, and assembled in Leapfrog Geothermal. This GDR submission includes all of the cleaned and reformatted (X (m), Y (m), elevation (m), processed data values) datasets used to build the Leapfrog Geodata model. Existing datasets were acquired from the GDR, from AltaRock, and from other sources. This yielded the following datasets: - Digital elevation model produced from LiDAR data by Ramsey and Bard, 2016 - MT surveys from 2006, 2011, 2014, and 2017 (including single inversions) - Gravity surveys from 2006, 2007, and 2011 (including single) - Earthquake catalogs from PNSN, LLNL, and the Newberry EGS Demonstration project - Seismic velocity model from Templeton et al., 2014 - The Frone, 2015 temperature model and a new one produced through extrapolating downhole temperature measurements and the SMU temperature at depth maps. Two versions of the new model are provided: 250 m spacing and 500 m spacing - EarthVision geologic model with alteration from Moser et al., 2016 - Well data from EGS well 55-29, deep geothermal wells, coreholes (GEO N-2 through 5) and several thermal gradient holes - "Newberry Well Data:" Location, simple lithology, directional survey data, and temperature data for the 34 wells and coreholes used in the Newberry PFA Although there are additional 2D datasets available in the area, such as aeromagnetic surveys, these were not included in the analysis. While it may be possible to project these datasets into three dimensions by assuming the surface measurements do not vary with depth, this method is associated with high uncertainty. Preexisting inversions of these data were unavailable, and inverting additional geophysical datasets is outside the scope of this project.

15 GEOTHERMAL ENERGY↗

ResStock Dataset 2024.1 Documentation

Public ResStock datasets provide credible, relevant, and accessible information on energy use and related non-energy metrics to a variety of stakeholders in the residential buildings space. The current public datasets include baseline building characteristics, timeseries (15-minute) energy consumption, and timeseries carbon emissions for the baseline (existing) U.S. housing stock and the U.S. housing stock with 10 "what-if" energy measure packages applied. This report documents a new public ResStock dataset to complement and build upon the existing public datasets. This dataset is specifically intended to be a resource for state and local decision-makers considering options for energy retrofits for their housing stock to reduce carbon emissions, energy use, and/or utility bills. These data consist of housing stock characteristics and modeled full-year energy consumption, carbon emission, energy bill, and energy burden data for the baseline U.S. housing stock as well as the U.S. housing stock with 260 "what-if" energy measure packages applied. These measure packages include measures related to the building envelope, appliances, pools and spas, lighting, water heating, and HVAC (including efficiency improvements and fuel switching with equipment at a range of performance levels) in a variety of combinations. This report provides methodology information on the generation of this dataset and serves as a key part of the dataset's public documentation.

buildings↗

A FAIR and AI-ready Higgs boson decay dataset

Abstract To enable the reusability of massive scientific datasets by humans and machines, researchers aim to adhere to the principles of findability, accessibility, interoperability, and reusability (FAIR) for data and artificial intelligence (AI) models. This article provides a domain-agnostic, step-by-step assessment guide to evaluate whether or not a given dataset meets these principles. We demonstrate how to use this guide to evaluate the FAIRness of an open simulated dataset produced by the CMS Collaboration at the CERN Large Hadron Collider. This dataset consists of Higgs boson decays and quark and gluon background, and is available through the CERN Open Data Portal. We use additional available tools to assess the FAIRness of this dataset, and incorporate feedback from members of the FAIR community to validate our results. This article is accompanied by a Jupyter notebook to visualize and explore this dataset. This study marks the first in a planned series of articles that will guide scientists in the creation of FAIR AI models and datasets in high energy particle physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗