Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The HydroBio Dataset: a new data resource for evaluating existing and potential hydropower capacity and freshwater biodiversity in the conterminous United States

Hydropower is a critical source of affordable and reliable electricity and energy system stability services in the United States. Opportunities to expand US hydropower production include retrofitting existing non-powered dams to produce power, retrofitting existing hydropower dams to improve efficiency or increase capacity, or constructing new hydropower infrastructure on currently unregulated river reaches. We created the HydroBio Dataset, which summarizes existing and potential hydropower capacity and freshwater biodiversity at the sub-basin scale in the conterminous US to contextualize existing and potential grid contributions with the freshwater ecosystems in which dams are situated. We demonstrate a use-case of this dataset by rescaling and comparing potential non-powered dam nominal capacity to rarity-threat-weighted freshwater species richness for sub-basins where both types of data exist. On average, normalized freshwater biodiversity exceeded normalized potential non-powered dam nominal capacity in these sub-basins. Potential non-powered dam nominal capacity was concentrated in sub-basins in the Upper Mississippi and Ohio hydrologic regions while freshwater biodiversity was concentrated in the South Atlantic-Gulf, Ohio, and Tennessee hydrologic regions. Additionally, non-powered dams and existing hydropower dams are located in sub-basins with similar indices of freshwater biodiversity. The HydroBio Dataset adds an additional ecological dimension of context to our understanding of current and potential future US hydropower capabilities and is a valuable decision support tool for stakeholders tasked with balancing gains in services to the US power grid with the public and environmental benefits of freshwater ecosystems.

Biodiversity↗

The Zooplankton International Geospatial dataset: A global repository of spatiotemporal freshwater zooplankton community composition data from lakes and reservoirs to support ecological research

Zooplankton transfer substantial energy in aquatic food webs and are used as indicators of environmental change. Syntheses of zooplankton community dynamics globally require datasets that span a wide range of environmental gradients; however, these datasets are limited due to methodological differences across programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we created the Zooplankton International Geospatial (ZIG) dataset, which includes original zooplankton, water physical and chemical variables, and lake morphometric data from 311 inland lakes and reservoirs. ZIG includes waterbodies ranging in size from 0.005 to 82,100 km2 and spanning broad latitudinal (−47.26 to 64.90) and longitudinal ranges (−165.04 to 176.53). Temporal coverage for individual waterbodies ranges between 1 and 60 yr with sampling frequency ranging from annually to weekly. With its extensive coverage and content, we consider ZIG to be a cornerstone for future investigations of global scale lake biodiversity change.

Figary, Stephanie [Cornell University, Ithaca, NY]↗

A traffic accident dataset for Chattanooga, Tennessee

This publication presents an annotated accident dataset which fuses traffic data from radar detection sensors, weather condition data, and light condition data with traffic accident data (as illustrated in Fig. 1) in a format that is easy to process using machine learning tools, databases, or data workflows. The purpose of this data is to analyze, predict, and detect traffic patterns when accidents occur. Each file contains a timeseries of traffic speeds, flows, and occupancies at the sensor nearest to the accident, as well as 5 neighboring sensors upstream and downstream. It also contains information about the accident type, date, and time. In addition to the accident data, we provide baseline data for typical traffic patterns during a given time of day. Overall, the dataset contains 6 months of annotated traffic data from November 2020 to April 2021. During this timeframe, and 361 accidents occurred in the monitored area around Chattanooga, Tennessee. This dataset served as the basis for a study on topology-aware automated accident detection for a companion publication [1].

97 MATHEMATICS AND COMPUTING↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

CryoDRGN-AI: neural ab initio reconstruction of challenging cryo-EM and cryo-ET datasets

Proteins and other biomolecules form dynamic macromolecular machines that are tightly orchestrated to move, bind, and perform chemistry. Cryo-electron microscopy (cryo-EM) and cryo-electron tomography (cryo-ET) can access the intrinsic heterogeneity of these complexes and are therefore key tools for understanding their function. However, 3D reconstruction of the collected imaging data presents a challenging computational problem, especially without any starting information, a setting termed ab initio reconstruction. Here, in this study, we introduce cryoDRGN-AI, a method leveraging an expressive neural representation and combining an exhaustive search strategy with gradient-based optimization to process challenging heterogeneous datasets. Using cryoDRGN-AI, we reveal new conformational states in large datasets, reconstruct previously unresolved motions from unfiltered datasets, and demonstrate ab initio reconstruction of biomolecular complexes from in situ data. With this expressive and scalable model for structure determination, we hope to unlock the full potential of cryo-EM and cryo-ET as a high-throughput tool for structural biology and discovery.

Levy, Axel [Stanford Univ., CA (United States); SL↗

Spatially distributed atmospheric boundary layer properties in Houston – A value-added observational dataset

Abstract In 2022, Houston, TX became a nexus for field campaigns aiming to further our understanding of the feedbacks between convective clouds, aerosols and atmospheric boundary layer (ABL) properties. Houston’s proximity to the Gulf of Mexico and Galveston Bay motivated the collection of spatially distributed observations to disentangle coastal and urban processes. This paper presents a value-added ABL dataset derived from observations collected by eight research teams over 46 days between 2 June - 18 September 2022. The dataset spans 14 sites distributed within a ~80-km radius around Houston. Measurements from three types of instruments are analyzed to objectively provide estimates of nine ABL parameters, both thermodynamic (potential temperature, and relative humidity profiles and thermodynamic ABL depth) and dynamic (horizontal wind speed and direction, mean vertical velocity, updraft and downdraft speed profiles, and dynamical ABL depth). Contextual information about cloud occurrence is also provided. The dataset is prepared on a uniform time-height grid of 1 h and 30 m resolution to facilitate its use as a benchmark for forthcoming numerical simulations and the fundamental study of atmospheric processes.

54 ENVIRONMENTAL SCIENCES↗

Accurate Dehydrogenation Enthalpies Dataset for Liquid Organic Hydrogen Carriers

This contribution presents a comprehensive extension of the QM9 dataset (originally at 133 K molecules) with the calculation of G4MP2 enthalpies for 9,841 molecules, featuring up to nine heavy atoms. We present QM9-LOHC, a (de)hydrogenation dataset of 10,373 reactions, including a minimum of 5.5% weight hydrogen storage capacity in line with the Department of Energy standards for Liquid Organic Hydrogen Carriers (LOHC). By utilizing the accurate quantum chemical method G4MP2 we expand the QM9 database and explore new avenues for the exploration of hydrogen storage technologies (electrochemical LOHCs, alkali metal-LOHCs, and mixtures of LOHCs). The QM9-LOHC dataset, with its focus on reactions that vary only by hydrogen saturation levels, provides a needed data resource for advancing the design and optimization of both conventional and innovative LOHC systems, and high-fidelity data for molecular discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PAH101: A GW+BSE Dataset of 101 Polycyclic Aromatic Hydrocarbon (PAH) Molecular Crystals

Abstract The excited-state properties of molecular crystals are important for applications in organic electronic devices. TheGWapproximation and Bethe-Salpeter equation (GW+BSE) is the state-of-the-art method for calculating the excited-state properties of crystalline solids with periodic boundary conditions. We present the PAH101 dataset ofGW+BSE calculations for 101 molecular crystals of polycyclic aromatic hydrocarbons (PAHs) with up to ~500 atoms in the unit cell. To the best of our knowledge, this is the firstGW+BSE dataset for molecular crystals. The data records include theGWquasiparticle band structure, the fundamental band gap, the static dielectric constant, the first singlet exciton energy (optical gap), the first triplet exciton energy, the dielectric function, and optical absorption spectra for light polarized along the three lattice vectors. The dataset can be used to (i) discover materials with desired electronic/optical properties, (ii) identify correlations between DFT andGW+BSE quantities, and (iii) train machine learned models to help in materials discovery efforts.

Science & Technology - Other Topics↗

Datasets of Faults in Variable Air Volume Terminal Units in a Multi-Zone Commercial Building

Faults in HVAC systems can decrease system efficiency and equipment lifespan, leading to 5%–30% of energy consumption being wasted in commercial buildings. We identified two common faults in HVAC variable air volume systems: a stuck damper fault in the variable air volume terminal unit and a discharge airflow sensor fault. We conducted three sets of damper stuck tests and two sets of airflow sensor tests, each including a fault-free scenario and scenarios with varying levels of faults, over one day. The faults were implemented in Oak Ridge National Laboratory’s two-story Flexible Research Platform building to generate a high-quality, well-controlled dataset covering fault-induced and fault-free scenarios. The test building, fault test scenarios, and data validation are described here. The open-source dataset includes 1 min intervals of weather and building data on the presence and absence of building faults. This dataset can be used to analyze the effects of HVAC system faults on system operation and indoor building conditions, and to develop or evaluate a fault detection and diagnosis algorithm.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dislocation-Grain Boundary Interaction Dataset for FCC Cu

Interactions between dislocations and grain boundaries play a major role in controlling the strength and ductility of structural materials. Experimentally, assessing and probing geometric and stress-based criteria at the local level for dislocation transmission through grain boundaries remains challenging. Therefore, there have been many efforts to systematically generate datasets of dislocation-grain boundary interactions (DGI) via computational models such as molecular dynamics simulations. So far, most DGI datasets have focused only on the subset of nominal minimum-energy grain boundary structures, which limits their applicability, especially to materials processed far from equilibrium. We present a comprehensive database of dislocation-grain boundary interactions for edge, screw, and 60° mixed dislocation with 330 <110> and 257 <112> symmetric tilt grain boundaries (total of 587) in FCC Cu consisting of 73 minimum-energy grain boundary structures and 514 metastable structures. The dataset contains the outcomes for 5234 unique interactions for various dislocation types, grain boundary structures, and applied shear stresses.

36 MATERIALS SCIENCE↗

Text-mined dataset of solid-state syntheses with impurity phases using Large Language Model

Solid-state synthesis is widely used to obtain various inorganic materials, such as battery materials and bulk thermoelectrics. Despite its prevalence, the process remains challenging due to the lack of a general theory and well-understood underlying reaction mechanisms. While prior works have successfully extracted structured datasets from literature, they often neglect product phase purity or yield. In this work, we construct a solid-state synthesis dataset consisting of 80,806 syntheses extracted with a large language model (LLM), including 18,869 reactions with impurity phase(s). Our dataset not only validates expected thermodynamic trends for impurity phase formation but also identifies challenging cases where impurity phases emerge even when the target phase is significantly more stable.

Lee, Sanghoon↗

Measurements of the temperature and $E$-mode polarization of the cosmic microwave background from the full 500-square-degree SPTpol dataset

Using the full four-year SPTpol 500 deg 2 dataset in both the 95 and 150 GHz frequency bands, we present measurements of the temperature and E-mode polarization of the cosmic microwave background (CMB), as well as the E-mode polarization autopower spectrum (EE) and temperature-E-mode cross-power spectrum (TE) in the angular multipole range 50 < ℓ < 8000. We find the SPTpol dataset to be self-consistent, passing several internal consistency tests based on maps, frequency bands, bandpowers, and cosmological parameters. The full SPTpol dataset is well-fit by the ΛCDM model, for which we find H 0 = 70.48 ± 2.16 km s -1 Mpc -1 and Ω m = 0.271 ± 0.026, when using only the SPTpol data and a Planck-based prior on the optical depth to reionization. The ΛCDM parameter constraints are consistent across the 95 GHz-only, 150 GHz-only, TE-only, and EE-only data splits. Between the ℓ < 1000 and ℓ > 1000 data splits, the ΛCDM parameter constraints are borderline consistent at the ∼2σ level. This consistency improves when including a parameter A L , the degree of lensing of the CMB inferred from the smearing of acoustic peaks. When marginalized over A L , the ΛCDM parameter constraints from SPTpol are consistent with those from Planck. In conclusion, the power spectra presented here are the most sensitive measurements of the lensed CMB damping tail to date for roughly ℓ > 1700 in TE and ℓ > 2000 in EE.

79 ASTRONOMY AND ASTROPHYSICS↗

Open Source Synergy: Developing and Validating PMU Data Analysis Techniques Using Open Source Tools and Datasets

This paper presents an exploration into the development and validation of data analysis approaches for Phasor Measurement Units (PMUs) using open-source datasets and tools. Various methods for event detection, event classification, frequency response, and oscillation analysis were tested. We leverage the capabilities of Archive Walker (AW), the Frequency Response Analysis Tool (FRAT), and the Oscillation Baselining and Analysis Tool (OBAT), all open-source tools, for efficient processing and analysis of synchrophasor data. The open-source Transmission Signature Library (TSL) dataset was employed as a dataset for a comprehensive evaluation to assess the performance and reliability of the proposed methods.

PMU, event analysis, oscillation, Frequency Respon↗

Chemical-specific Parameters Dataset

The chemical-specific parameters dataset is searchable for physicochemical information for multiple chemicals simultaneously. After selecting chemicals of interest and the desired parameters, the RAIS will generate a table containing the values, chosen according to an established hierarchy. Results can be downloaded in Excel format. Over 40 parameters are available, including melting point, boiling point, density, density, vapor pressure, water solubility, and Henry’s Law constants. Thirteen primary sources are used to populate the dataset of chemical-specific parameters. These values should be used in cancer risk and noncancer hazard assessments for the calculation of preliminary remediation goals (PRGs), hazard characterization, and transport modeling. Users can select up to 1000 chemicals per query. The dataset supports environmental risk assessments, regulatory decision-making, and environmental planning with tools for benchmarking against risk-based standards. This structured approach ensures a robust evaluation of environmental risks tailored to regulatory needs.

Dolislager, Fred [Oak Ridge National Laboratory (O↗

Radionuclide-specific Parameters Dataset

The radionuclide-specific parameters dataset is searchable for radiological information for multiple isotopes simultaneously. After selecting radionuclides of interest and the desired parameters, the RAIS will generate a table containing the values, chosen according to an established hierarchy. Results can be downloaded in Excel format. 50 parameters are available, including atomic number, soil to animal transfer coefficients, plant uptake coefficients, half-life, specific activity, and water solubility. Seven primary sources are used to populate the dataset of radiological-specific parameters. These values should be used in cancer risk assessments for the calculation of preliminary remediation goals (PRGs), hazard characterization, and transport modeling. Users can select up to 1000 radionuclides per query. The dataset supports environmental risk assessments, regulatory decision-making, and environmental planning with tools for benchmarking against risk-based standards. This structured approach ensures a robust evaluation of environmental risks tailored to regulatory needs.

Manning, Karessa [Oak Ridge National Laboratory (O↗

Road Lidar Dataset for the TxDOT Austin District

This is a road lidar data collection for developing road elevation models and road inundation mapping methodologies, a joint work between ORNL and The University of Texas at Austin. This dataset is generated as part of the flood transportation infrastructure, partly funded by the NOAA CIROH project. ORNL is a project partner for high-performance computing-empowered flood inundation mapping methodology R&D. The dataset is computed using a GPU-accelerated lidar data processing workflow developed at ORNL. The lidar data source is from TxGIO, the state lidar data collection site. The output dataset is in two formats: laz and copc. It is organized by TxDOT's maintenance sections, covering the Austin District. Data size: 3.86 billion road lidar points, 1.67% of the entire lidar data input Projection: EPSG:32614 (WGS84/UTM zone 14N) Website: https://web.corral.tacc.utexas.edu/nfiedata/road3d/austin_district/AustinMaintenanceSections_H_epsg6343_V_epsg5703/ LICENSE FOR USE -- MAPS AND DATA DISCLAIMER This resource is shared under the Creative Commons Attribution CC BY, http://creativecommons.org/licenses/by/4.0/ MAPS AND DATA DISCLAIMER The Oak Ridge National Laboratory (ORNL) shall not be held liable for improper or incorrect use of the data described or information contained on this map or associated series of maps. The data and related map graphics are not legal, land survey or engineering documents and are not intended to be used as such. ORNL gives no warranty, express or implied, as to the accuracy, reliability, utility or completeness of this information. The user of these maps and data assumes all responsibility and risk for the use of the maps and data. ORNL disclaims all warranties, representations or endorsements either express or implied, with regard to the information contained in this map product, including, but not limited to, all implied warranties of merchantability, fitness for a particular purpose or non-infringement. This preliminary map product is for research and review purposes only. It is not intended to be used for emergency management operational or life safety decisions at the local or regional governmental level or by the general public. Users requiring information regarding hazardous conditions or meteorological conditions for specific geographic areas should consult directly with their city or county emergency management office.

54 ENVIRONMENTAL SCIENCES↗

Driver Identification Dataset

The ORNL Driver Identification Dataset was created to collect and analyze driving behavior data from 50 different drivers. Each driver operated a 2014 Kenworth T270 Class 6 truck around Fort Collins, Colorado while various data sources recorded their driving behavior and vehicle performance. The dataset includes CANbus (Controller Area Network) data, GPS data, inertial measurement data, and biometric data from a heart rate monitor. A cyberattack was executed during each drive, which caused multiple dashboard warning lights to illuminate and set the tachometer and speedometer to zero, regardless of actual speed. The attack was stopped either after one minute or if the driver pulled over. By downloading the dataset, you agree to the following: 1) I will not use or disclose the data for any purpose other than Research as that term is defined in 10 CFR 745.102. 2) I will not, under any circumstances, request or accept private or linking identifiers for the data used. 3) I will not attempt to determine the identity of the individuals associated with the data. 4) I will use appropriate safeguards to prevent the use or disclose of the data for any purpose other than Research.

99 GENERAL AND MISCELLANEOUS↗

TxDOT Road Elevation Model Dataset

This dataset provides three formats of Road Elevation Model (REM) data: 3D road line/polygon GeoPackage (GPKG), road lidar LAZ and COPC LAZ, and road digital surface model (DSM) GeoTIFF. Data are produced from the ~50TB TxGIO (formerly TNRIS) state lidar collections. This dataset is currently organized by maintenance section in each TxDOT district. Computation is done on GPU computing resources at Oak Ridge National Laboratory (ORNL), through a Strategic Partnership Project with UT Austin and an NSF ACCESS computing allocation award that enables fast massive data movement between TACC Corral and ORNL CADES/OLCF using Globus. In addition to this release from ORNL, a copy of this dataset can also be downloaded at https://web.corral.tacc.utexas.edu/nfiedata/road3d/.

13 HYDRO ENERGY↗