Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION↗

2024 Buildings Technology Baseline: Dataset Documentation

The Buildings Technology Baseline is a curated and regularly updated dataset of current and projected performance, retail, and installed price data for all major building energy technologies needed to enable cost/benefit analyses. Building technology analyses require an up-to-date understanding of installation costs and cost-effectiveness of key building energy efficiency technologies. The dataset was assembled by Guidehouse during fiscal year 2024. Data was gathered from the 2024 National Residential Efficiency Measures Database (NREMDB), the 2023 Energy Information Administration Updated Buildings Sector Appliance and Equipment Costs and Efficiencies ("EIA Building Data Report"), DOE Lighting Market Model, the 2023 RSMeans database, and the 2020 Grid-Interactive Efficient Building Technology Cost, Performance, and Lifetime Characteristics ("GEB Data Report"), Lawrence Berkeley National Laboratory, various literature, as well as new data from online retailers, stakeholder interviews, and contractor databases in 2023 and 2024. The dataset has been reviewed by subject matter experts at NREL and DOE. The 2024 dataset release is intended to be a starting point for interested users to provide feedback. This database is not intended to provide specific cost estimates for a specific project. The cost estimates do not include any rebates or tax incentives that may be available for the measures. Rather, it is meant to help determine which measures may be more cost-effective. The National Renewable Energy Laboratory (NREL) makes every effort to ensure accuracy of the data; however, NREL does not assume any legal liability or responsibility for the accuracy or completeness of the information.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Development of A High-Resolution Dataset for Solar Resource Adequacy Studies

High-resolution, long-term solar dataset is essential for characterizing the variability of solar energy resources and for informing strategies that ensure grid reliability and resilience in grid systems with high levels of solar energy integration. We present the development of a new 4-km, hourly Earth system dataset for the contiguous United States (CONUS), using a statistical downscaling approach that integrates the National Solar Radiation Database (NSRDB) with regional Earth system model projections. The new high-resolution Earth system dataset includes key variables - GHI, DNI, DHI, surface air temperature, and wind speed - under two future scenarios. Preliminary results show a reasonable agreement with NSRDB observations, with nBias less than 1% for GHI across CONUS. The dataset is expected to support in-depth analyses of extreme weather impacts and provide input to resource adequacy for future energy systems with diverse generation sources.

14 SOLAR ENERGY↗

BUTTER - Empirical Deep Learning Dataset

The BUTTER Empirical Deep Learning Dataset represents an empirical study of the deep learning phenomena on dense fully connected networks, scanning across thirteen datasets, eight network shapes, fourteen depths, twenty-three network sizes (number of trainable parameters), four learning rates, six minibatch sizes, four levels of label noise, and fourteen levels of L1 and L2 regularization each. Multiple repetitions (typically 30, sometimes 10) of each combination of hyperparameters were preformed, and statistics including training and test loss (using a 80% / 20% shuffled train-test split) are recorded at the end of each training epoch. In total, this dataset covers 178 thousand distinct hyperparameter settings ("experiments"), 3.55 million individual training runs (an average of 20 repetitions of each experiments), and a total of 13.3 billion training epochs (three thousand epochs were covered by most runs). Accumulating this dataset consumed 5,448.4 CPU core-years, 17.8 GPU-years, and 111.2 node-years.

Array↗

County-Level Hourly Renewable Capacity Factor Dataset for the ReEDS Model

This dataset contains hourly capacity factors for each renewable resource class and region (in this case, county). Technologies like large-scale utility PV (UPV), onshore wind, offshore wind, and concentrating solar power (CSP) are included. The dataset contains 7 years of hourly weather data (2007-2013) for different sites across the US and is used as one of the inputs to the ReEDS-2.0 model (see the "ReEDS 2.0 GitHub Repository" resource link below), developed by NREL. The weather profiles apply to any capacity that exists or is built in each region and class. This helps calculate the generation that can be provided using these resources. Open, reference, and limited are 3 scenarios based on land-use allowance, derived from the Renewable Energy Potential (reV) model developed by NREL, which helps generate supply curves for renewable technologies and assess the maximum potential of renewable resources in a designated area. Each zipped file in this dataset corresponds to a technology and contains the respective land-use scenario files required to run that technology in ReEDS. To use this dataset, download and place the extracted files in the locally cloned ReEDS repository inside one of the folders (inputs/variability/multi_year). After completing this copy, upon running the ReEDS model at the county-level spatial resolution for respective analysis purposes, the program will detect the presence of these files and will not fail.

Array↗

A computational pipeline to generate a synthetic dataset of metal ion sorption to oxides for AI/ML exploration

The charged mineral/electrolyte interfaces are ubiquitous in the surface and subsurface–including the surroundings of the geological disposal sites for radioactive waste. Therefore, understanding how ions interact with charged surfaces is critically important for predicting radionuclide mobility in the case of waste leakage. At present, the Surface Complexation Models (SCMs) are the most successful thermodynamic frameworks to describe ion retention by mineral surfaces. SCMs are interfacial speciation models that account for the effect of the electric field generated by charged surfaces on sorption equilibria. These models have been successfully used to analyze and interpret a broad range of experimental observations including potentiometric and electrokinetic titrations or spectroscopy. Unfortunately, many of the current procedures to solve and fit SCM to experimental data are not optimal, which leads to a non-transferable or non-unique description of interfacial electrostatics and consequently of the strength and extent of ion retention by mineral surfaces. Recent developments in Artificial Intelligence (AI) offer a new avenue to replace SCM solvers and fitting algorithms with trained AI surrogates. Unfortunately, there is a lack of a standardized dataset covering a wide range of SCM parameter values available for AI exploration and training–a gap filled by this study. Here, we described the computational pipeline to generate synthetic SCM data and discussed approaches to transform this dataset into AI-learnable input. First, we used this pipeline to generate a synthetic dataset of electrostatic properties for a broad range of the prototypical oxide/electrolyte interfaces. The next step is to extend this dataset to include complex radionuclide sorption and complexation, and finally, to provide trained AI architectures able to infer SCMs parameter values rapidly from experimental data. Here, we illustrated the AI-surrogate development using the ensemble learning algorithms, such as Random Forest and Gradient Boosting. These surrogate models allow a rapid prediction of the SCM model parameters, do not rely on an initial guess, and guarantee convergence in all cases.

Li, Chunhui↗

Machine learning prediction of the mechanical properties of refractory multicomponent alloys based on a dataset of phase and first principles simulation

In this work, a dataset including structural and mechanical properties of refractory multicomponent alloys was developed by fusing computations of phase diagram (CALPHAD) and density functional theory (DFT). The refractory multicomponent alloys, also named refractory complex concentrated alloys (CCAs) which contain 2–5 types of refractory elements were constructed based on Special Quasi-random Structure (SQS). The phase of alloys was predicted using CALPHAD and the mechanical property of alloys with stable and single body-centered cubic (BCC) at high temperature (over 1,500°C) was investigated using DFT-based simulation. As a result, a dataset with 393 refractory alloys and 12 features, including volume, melting temperature, density, energy, elastic constants, mechanical moduli, and hardness, were produced. To test the capability of the dataset on supporting machine learning (ML) study to investigate the property of CCAs, CALPHAD, and DFT calculations were compared with principal components analysis (PCA) technique and rule of mixture (ROM), respectively. It is demonstrated that the CALPHAD and DFT results are more in line with experimental observations for the alloy phase, structural and mechanical properties. Furthermore, the data were utilized to train a verity of ML models to predict the performance of certain CCAs with advanced mechanical properties, highlighting the usefulness of the dataset for ML technique on CCA property prediction.

36 MATERIALS SCIENCE↗

metabCombiner 2.0: Disparate Multi-Dataset Feature Alignment for LC-MS Metabolomics

Liquid chromatography–high-resolution mass spectrometry (LC-HRMS), as applied to untargeted metabolomics, enables the simultaneous detection of thousands of small molecules, generating complex datasets. Alignment is a crucial step in data processing pipelines, whereby LC-MS features derived from common ions are assembled into a unified matrix amenable to further analysis. Variability in the analytical factors that influence liquid chromatography separations complicates data alignment. This is prominent when aligning data acquired in different laboratories, generated using non-identical instruments, or between batches from large-scale studies. Previously, we developed metabCombiner for aligning disparately acquired LC-MS metabolomics datasets. Here, we report significant upgrades to metabCombiner that enable the stepwise alignment of multiple untargeted LC-MS metabolomics datasets, facilitating inter-laboratory reproducibility studies. To accomplish this, a “primary” feature list is used as a template for matching compounds in “target” feature lists. We demonstrate this workflow by aligning four lipidomics datasets from core laboratories generated using each institution’s in-house LC-MS instrumentation and methods. We also introduce batchCombine, an application of the metabCombiner framework for aligning experiments composed of multiple batches. metabCombiner is available as an R package on Github and Bioconductor, along with a new online version implemented as an R Shiny App.

97 MATHEMATICS AND COMPUTING↗

FLUXNET-CH 4 : a global, multi-ecosystem dataset and analysis of methane seasonality from freshwater wetlands

Methane (CH 4 ) emissions from natural landscapes constitute roughly half of global CH4 contributions to the atmosphere, yet large uncertainties remain in the absolute magnitude and the seasonality of emission quantities and drivers. Eddy covariance (EC) measurements of CH4 flux are ideal for constraining ecosystem-scale CH 4 emissions due to quasi-continuous and high-temporal-resolution CH 4 flux measurements, coincident carbon dioxide, water, and energy flux measurements, lack of ecosystem disturbance, and increased availability of datasets over the last decade. Here, we (1) describe the newly published dataset, FLUXNET-CH 4 Version 1.0, the first open-source global dataset of CH 4 EC measurements (available at https://fluxnet.org/data/fluxnet-ch4-community-product/, last access: 7 April 2021). FLUXNET-CH 4 includes half-hourly and daily gap-filled and non-gap-filled aggregated CH4 fluxes and meteorological data from 79 sites globally: 42 freshwater wetlands, 6 brackish and saline wetlands, 7 formerly drained ecosystems, 7 rice paddy sites, 2 lakes, and 15 uplands. Then, we (2) evaluate FLUXNET-CH 4 representativeness for freshwater wetland coverage globally because the majority of sites in FLUXNET-CH4 Version 1.0 are freshwater wetlands which are a substantial source of total atmospheric CH 4 emissions; and (3) we provide the first global estimates of the seasonal variability and seasonality predictors of freshwater wetland CH4 fluxes. Our representativeness analysis suggests that the freshwater wetland sites in the dataset cover global wetland bioclimatic attributes (encompassing energy, moisture, and vegetation-related parameters) in arctic, boreal, and temperate regions but only sparsely cover humid tropical regions. Seasonality metrics of wetland CH 4 emissions vary considerably across latitudinal bands. In freshwater wetlands (except those between 20°S to 20°N) the spring onset of elevated CH 4 emissions starts 3 d earlier, and the CH 4 emission season lasts 4 d longer, for each degree Celsius increase in mean annual air temperature. On average, the spring onset of increasing CH 4 emissions lags behind soil warming by 1 month, with very few sites experiencing increased CH 4 emissions prior to the onset of soil warming. In contrast, roughly half of these sites experience the spring onset of rising CH 4 emissions prior to the spring increase in gross primary productivity (GPP). The timing of peak summer CH 4 emissions does not correlate with the timing for either peak summer temperature or peak GPP. Our results provide seasonality parameters for CH 4 modeling and highlight seasonality metrics that cannot be predicted by temperature or GPP (i.e., seasonality of CH 4 peak). FLUXNET-CH 4 is a powerful new resource for diagnosing and understanding the role of terrestrial ecosystems and climate drivers in the global CH 4 cycle, and future additions of sites in tropical ecosystems and site years of data collection will provide added value to this database. All seasonality parameters are available at https://doi.org/10.5281/zenodo.4672601 (Delwiche et al., 2021). Additionally, raw FLUXNET-CH 4 data used to extract seasonality parameters can be downloaded from https://fluxnet.org/data/fluxnet-ch4-community-product/ (last access: 7 April 2021), and a complete list of the 79 individual site data DOIs is provided in Table 2 of this paper.

54 ENVIRONMENTAL SCIENCES↗

Algorithmically detected rain-on-snow flood events in different climate datasets: a case study of the Susquehanna River basin

Abstract. Rain-on-snow (RoS) events in regions of ephemeral snowpack – such as the northeastern United States – can be key drivers of cool-season flooding. We describe an automated algorithm for detecting basin-scale RoS events in gridded climate data by generating an area-averaged time series and then searching for periods of concurrent precipitation, surface runoff, and snowmelt exceeding predefined thresholds. When evaluated using historical data over the Susquehanna River basin (SRB), the technique credibly finds RoS events in published literature and flags events that are followed by anomalously high streamflow as measured by gauge data along the river. When comparing four different datasets representing the same 21-year period, we find large differences in RoS event magnitude and frequency, primarily driven by differences in estimated surface runoff and snowmelt. Using dataset-specific thresholds improves agreement between datasets but does not account for all discrepancies. We show that factors such as meteorological forcing and coupling frequency, as well as choice of land surface model, play roles in how data products capture these compound extremes and suggest care is to be taken when climate datasets are used by stakeholders for operational decision-making.

54 ENVIRONMENTAL SCIENCES↗

Farm Practice Typologies as a Strategy for Management-Relevant Land Use and Land Cover Mapping in the Great Lakes Region (Version 1) [Dataset]

Dataset overview and development This dataset provides spatially explicit agricultural land-use and land-management typologies developed for the Great Lakes Region (GLR) at the farm-parcel level. The typologies were designed to characterize not only the land-use and land-cover (LULC) associated with individual agricultural farm parcels, but also the land-management practices (LMPs), including irrigation, tile drainage, and conservation easements, occurring within those parcels and how these characteristics change through time. The dataset contains four related typology products: Annual integrated typology – describes the combined LULC and land-management characteristics for each farm parcel for individual years. LULC transition typology – describes the temporal pattern of LULC change for each farm parcel across the study period (2008-2023). LMP trend typology – describes the temporal pattern in the occurrence of LMPs for each farm parcel across the study period. Multi-year integrated typology – combines the LULC transition typology and LMP trend typology to provide an integrated characterization of long-term land-use and management patterns. Purpose of the dataset The purpose of these products is to provide a management-relevant integrated and consistent framework for evaluating the spatial and temporal organization of agricultural landscapes across the GLR. The resulting typologies can: support landscape-scale environmental and land-use analysis; provide spatial information relevant to land-management strategies, conservation planning, policy development, and program evaluation; characterize spatial patterns of agricultural land use and management; examine changes in agricultural landscapes through time; and identify persistent, transitional, and changing agricultural systems. Please refer to the README file provided in Files for more details.

Agriculture↗

LCLS RF Station Phase Anomaly Candidate Dataset

A public anomaly detection dataset constructed from RF station faults for phase at SLAC's LCLS (Linac Coherent Light Source). We have compiled a dataset of the RF station diagnostic phase data and the beam-position monitor (BPM) signals, alongside the hand labels, for a labeled study period. The dataset consists of two HDF5 files (one for train and one for test) containing the raw data, two CSV files containing information about the candidates. The CSV file for the test dataset also contains the label.

Liang, Jia [Stanford Univ., CA (United States). In↗

Legacy Survey of Space and Time Data Preview 1: raw dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the raw dataset type. These are unprocessed images from the LSST Commissioning Camera. This release contains 16,125 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: visit_image dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the visit_image dataset type. These are individual processed and calibrated sky images obtained from a single observation with a single filter. This release contains 15,972 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: difference_image dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the difference_image dataset type. These are images created by subtracting a template image from a visit image. This release contains 15,972 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: deep_coadd dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the deep_coadd dataset type. These are the combination of multiple processed, calibrated, and background- subtracted images, for a patch of sky, for each of the six filters. This release contains 2,644 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: template_coadd dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the template_coadd dataset type. These are the combination of processed images with the best seeing, for a patch of sky and for each of the six LSST filters. Used to create difference images. This release contains 2,730 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: survey property dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the survey property dataset type. These are healSparse property maps for the survey. This release contains 84 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗