Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

MArVD2: a machine learning enhanced tool to discriminate between archaeal and bacterial viruses in viral datasets

Abstract Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.

Vik, Dean (ORCID:000000027546899X)↗

Highly transferable atomistic machine-learning potentials from curated and compact datasets across the periodic table

Machine learning atomistic potentials trained using density functional theory (DFT) datasets allow for the modeling of complex material properties with near-DFT accuracy while imposing a fraction of its computational cost. The curation of the DFT datasets can be extensive in size and time-consuming to train and refine. In this study, we focus on addressing these barriers by developing minimalistic and flexible datasets for many elements in the periodic table regardless of their mass, electronic configuration, and ground state lattice. These DFT datasets have on average, ~4000 different structures and 27 atoms per structure, which we found sufficient to maintain the predictive accuracy of DFT properties and notably with high transferability. We envision these highly curated training sets as starting points for the community to expand, modify, or use with other machine learning atomistic potential models, whatever may suit individual needs, further accelerating the utilization of machine learning as a tool for material design and discovery.

42 ENGINEERING↗

Exploring the growth index γ L : Insights from different CMB dataset combinations and approaches

In this study we investigate the growth index γ L , which characterizes the growth of linear matter perturbations, while analysing different cosmological datasets. We compare the approaches implemented by two different patches of the cosmological solver : and __. In our analysis we uncover a deviation of the growth index from its expected Λ CDM value of 0.55 when utilizing the Planck dataset, both in the case and in the __ case, but in opposite directions. This deviation is accompanied by a change in the direction of correlations with derived cosmological parameters. However, the incorporation of cosmic microwave background lensing data helps reconcile γ L with its Λ -cold dark matter value in both cases. Conversely, the alternative ground-based telescopes Atacama Cosmology Telescope and South Pole Telescope consistently yield growth index values in agreement with γ L = 0.55 . We conclude that the presence of the A lens problem in the Planck dataset contributes to the observed deviations, underscoring the importance of additional datasets in resolving these discrepancies. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Collection: TD-DFT and EOM-CCSD Calculations for the GDB-9-Ex Dataset

We present two datasets that contain quantum chemical electronic structure calculations for organic molecules from the GDB-9-Ex dataset. The “GDB-9-Ex_TD-DFT-PBE0” dataset contains calculations performed using the time-dependent density functional theory (TD-DFT) first principles method, and the “GDB-9-Ex_EOMCCSD” dataset contains calculations performed using the equation-of-motion coupled cluster (EOM-CCSD) method. Both types of calculations were performed using the ORCA software and provided ultraviolet-visible spectra with a high level of accuracy.

Mehta, Kshitij [Oak Ridge National Laboratory (ORN↗

Towards Diverse and Representative Global Pretraining Datasets for Remote Sensing Foundation Models

The design of a pretraining dataset is emerging as a critical component for the generality of foundation models. In the remote sensing realm, large volumes of imagery and benchmark datasets exist that can be leveraged to pretrain foundation models, however using this imagery in absence of a well-crafted sampling strategy is inefficient and has the potential to create biased and less generalizable models. Here, we provide a discussion and vision for the curation and assessment of pretraining datasets for remote sensing geospatial foundation models. We highlight the importance of geographic, temporal, and image acquisition diversity and review possible strategies to enable such diversity at global scale. In addition to these characteristics, support for various spatial-temporal pretext tasks within the dataset is also critical. Ultimately, our primary objective is to place emphasis on and draw attention to the data curation stage of the foundation model development pipeline. By doing so, we think it is possible to reduce biases of geospatial foundation models, as well as enable broader generalization to downstream remote sensing tasks and applications.

Arndt, Jacob↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗

RADAI: A Large-Scale Realistic Dataset for Radiation Detection Algorithm Development

Open, realistic datasets are essential for developing and benchmarking radiation detection algorithms, yet they remain scarce. The Radiological Anomaly Detection and Identification (RADAI) project was develop to create datasets that meet the training and testing needs for sophisticated radiation detection algorithms. The RADAI dataset is a large-scale synthetic resource that integrates high-fidelity Monte Carlo simulations with realistic urban scenarios to capture both background variability and source signatures. RADAI models construction-material NORM, people and vehicles, urban clutter, and dynamic environmental effects such as cosmic-ray and rain-induced transients, and they provide list-mode detector data with motion and response modeling suitable for algorithm training and evaluation. The RADAI project resulted in three publicly-released complementary datasets together with an online scoring portal for standardized performance assessment and an open software toolkit that supports data access, augmentation, model development, and evaluation. These resources enable reproducible comparisons across methods and promote rigorous studies at the scale required by contemporary machine learning. By grounding algorithm development in realistic, well-documented conditions, RADAI supports progress toward more robust detection, identification, and localization in complex urban environments.

Ghawaly, James M. [Division of Computer Science an↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

BAMCensus (The Behavior and Advanced Mobility Census Dataset Aggregator) [SWR-25-120]

This software is a high-performance tool developed in Rust for downloading and processing large-scale geospatial datasets, specifically focusing on US Census data. It is designed to address scaling limitations found in existing tools, such as R's [tidycensus](https://walker-data.com/tidycensus/), by providing performant streaming dataset JOIN operations between various US Census datasets (like ACS and LEHD) and their corresponding geometries stored on the TIGER/Lines web server. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. Its primary motivation stems from the need for a high-performance solution to combine spatial datasets with graph traversals within the context of mobility analysis tooling being developed at NREL's Behavior and Advanced Mobility (BAM) group.

Fitzgerald, Robert [National Renewable Energy Labo↗

Strictly Enforcing Invertibility and Conservation in CNN-Based Super Resolution for Scientific Datasets

Abstract Recently, deep convolutional neural networks (CNNs) have revolutionized image “super resolution” (SR), dramatically outperforming past methods for enhancing image resolution. They could be a boon for the many scientific fields that involve imaging or any regularly gridded datasets: satellite remote sensing, radar meteorology, medical imaging, numerical modeling, and so on. Unfortunately, while SR-CNNs produce visually compelling results, they do not necessarily conserve physical quantities between their low-resolution inputs and high-resolution outputs when applied to scientific datasets. Here, a method for “downsampling enforcement” in SR-CNNs is proposed. A differentiable operator is derived that, when applied as the final transfer function of a CNN, ensures the high-resolution outputs exactly reproduce the low-resolution inputs under 2D-average downsampling while improving performance of the SR schemes. The method is demonstrated across seven modern CNN-based SR schemes on several benchmark image datasets, and applications to weather radar, satellite imager, and climate model data are shown. The approach improves training time and performance while ensuring physical consistency between the super-resolved and low-resolution data. Significance Statement Recent advancements in using deep learning to increase the resolution of images have substantial potential across the many scientific fields that use images and image-like data. Most image super-resolution research has focused on the visual quality of outputs, however, and is not necessarily well suited for use with scientific data where known physics constraints may need to be enforced. Here, we introduce a method to modify existing deep neural network architectures so that they strictly conserve physical quantities in the input field when “super resolving” scientific data and find that the method can improve performance across a wide range of datasets and neural networks. Integration of known physics and adherence to established physical constraints into deep neural networks will be a critical step before their potential can be fully realized in the physical sciences.

54 ENVIRONMENTAL SCIENCES↗

DEEPEN 3D PFA Index Models for Exploration Datasets at Newberry Volcano

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), index models needed to be developed to map values in geoscientific exploration datasets to favorability index values. This GDR submission includes those index models. Index models were created by binning values in exploration datasets into chunks based on their favorability, and then applying a number between 0 and 5 to each chunk, where 0 represents very unfavorable data values and 5 represents very favorable data values. To account for differences in how exploration methods are used to detect each play component, separate index models are produced for each exploration method for each component of each play type. Index models were created using histograms of the distributions of each exploration dataset in combination with literature and input from experts about what combinations of geophysical, geological, and geochemical signatures are considered favorable at Newberry. This is in attempt to create similar sized bins based on the current understanding of how different anomalies map to favorable areas for the different types of geothermal plays (i.e., conventional hydrothermal, superhot EGS, and supercritical). For example, an area of partial melt would likely appear as an area of low density, high conductivity, low vp, and high vp/vs. This means that these target anomalies would be given high (4 or 5) index values for the purpose of imaging the heat source. To account for differences in how exploration methods are used to detect each play component, separate index models are produced for each exploration method for each component of each play type. Index models were produced for the following datasets: - Geologic model - Alteration model - vp/vs - vp - vs - Temperature model - Seismicity (density*magnitude) - Density - Resistivity - Fault distance - Earthquake cutoff depth model

15 GEOTHERMAL ENERGY↗

The LAKE model input dataset for three Arctic lakes

This dataset contains meteorological data collected for three Arctic lakes and compiled to satisfy input requirements of the LAKE 2.0 model. The dataset was generated to act as a benchmarking dataset for future model-data inter-comparisons. The LAKE 2.0 model simulates temperatures within the water later and the sedimentary layer of a lake. The LAKE2.0. is an open-source code and available to download via this weblike http://tesla.parallel.ru/Viktor/LAKE/-/wikis/LAKE-model (last visit July 14, 2021). The meteorological data are required to simulate the surface energy balance at the surface of a lake. This dataset includes a compilation of the meteorological data pulled from multiple data streams, including National Oceanic and Atmospheric Administration (NOAA) climate data, Circumarctic Lakes Observation Network (CALON) data, and the United States Geological Survey (USGS) data. The data were compiled for three Arctic lakes: FoxDen (66.55877, -164.45670), Atqasuk (70.452497, -156.951984), and Toolik (68.63150, -149.60740). Each meteorological data is in comma-delimited format (file extension ‘.dat’) and includes eight columns: Temperature [K], Pressure [Pa], longwave downward radiation [W/m2], shortwave downward radiation [W/m2], “U” wind speed [m/s], ”V” wind speed [m/s], humidity [kg/kg], precipitation [m/s]. In addition to the meteorological data file, we included setup and driver files. The Toolik lake is the deepest out of three lakes and has inflowing and outflowing groundwater data. InflowOutflowREADME.txt has more information about inflow and outflow flies. The other two lakes are much shallower and modeled as a closed system (i.e. no water inflow or outflow).

54 ENVIRONMENTAL SCIENCES↗

Analysis and Integration of the Hydraulic Fracturing Test Site-2 (HFTS-2) Comprehensive Dataset

Hydraulic Fracturing Test Site-2 (HFTS-2) is a field-based research experiment performed in the Permian (Delaware) Basin. The unique aspect of this program was the acquisition of a unique, comprehensive, diagnostic dataset. Additionally, shorter parent wells drilled three years before the child wells offered clear distinction between the stages influenced by parent-child effects and the stages without any effects. The goal of this study was to analyze and integrate this comprehensive diagnostic dataset to understand the areal and vertical extent of hydraulic fractures (HF). The paper also provides insights on the effects of parent wells’ depletion on child well HF geometry based on various monitoring methods and subsurface models. Areal and vertical coverage for all HFTS-2 wells during stimulation and depletion was estimated based on analysis and interpretation of diagnostics and advanced modeling results. HFTS-2 diagnostics included microseismic (MS), pre- and post-stimulation logs and cores, bottomhole gauges, and fiber optic (FO) data. The diagnostics results (MS, FO, image logs) were integrated and used to calibrate subsurface models. Additional field tests were designed and implemented for depletion monitoring. The tailored program for monitoring depletion included vertical and slant well pressures, interference testing, and a vertical strain depletion trial. Areal Coverage: Conventional MS (and FO MS) were used to compute HF dimensions, which were compared with diagnostics (FO strain, gauge, image logs) observations and calibrated subsurface models. A post-production interference test did not show offset well communication. Vertical Coverage: Vertical coverage during stimulation was monitored using a vertical monitoring well. The stronger mechanical strain signals showed good correlation with MS event intensities, geomechanical properties, and gauge inferences. Vertical depletion was estimated based on vertical/slant well gauges and strain depletion tests. Parent-Child Effects: Diagnostics and calibrated subsurface models show asymmetry in child well fracture geometries for stages that overlap parent wells. Child well image logs serve as a good indicator for parent Downloaded from http://onepetro.org/URTECONF/proceedings-pdf/21URTC/2-21URTC/D021S031R004/2477423/urtec-2021-5241-ms.pdf/1 by Carol Worster on 28 February 2022 URTeC 5241 well HF tracking. Child well MS events had an eastward bias, in line with pre-stimulation image logs, and was confirmed by parent well frac hits. Novel/Additive Information: The dataset presents a unique, over-constrained problem space to compare independent techniques to arrive at HF metrics (i.e., stimulation height and/or half-length), unlike a single source dataset, in which calibration is done using available data to guide predictions. Here, the asymmetry in HF geometry seen in the stages influenced by parent-child effects offers unique insights into well spacing and landing, which are key capital decisions the unconventional resources industry is seeking to optimize.

58 GEOSCIENCES↗

PubDAS: A PUBlic Distributed Acoustic Sensing Datasets Repository for Geosciences

During the past few years, distributed acoustic sensing (DAS) has become an invaluable tool for recording high-fidelity seismic wavefields with great spatiotemporal resolutions. However, the considerable amount of data generated during DAS experiments limits their distribution with the broader scientific community. Such a bottleneck inherently slows down the pursuit of new scientific discoveries in geosciences. Here, we introduce PubDAS—the first large-scale open-source repository where several DAS datasets from multiple experiments are publicly shared. PubDAS currently hosts eight datasets covering a variety of geological settings (e.g., urban centers, underground mines, and seafloor), spanning from several days to several years, offering both continuous and triggered active source recordings, and totaling up to ~90 TB of data. Here this article describes these datasets, their metadata, and how to access and download them. Some of these datasets have only been shallowly explored, leaving the door open for new discoveries in Earth sciences and beyond.

54 ENVIRONMENTAL SCIENCES↗

Modeling Air Handling Units to Create a Diverse Fault Dataset for FDD Innovation: Lessons Learned and Recommendations

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault datasets for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling unit and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for the air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, two detailed AHU models, which included the single duct AHU and dual duct AHU developed in the Modelica language and HVACSIM+ were employed to carry out annual simulations of numerous common sensor faults, mechanical faults, and control sequence faults. The fault inclusive data were then validated by comparing fault effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. We report some lessons learnt during the efforts of validating the high volumes of the FDD data sets. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

Curation of Ground-Truth Validated Benchmarking Datasets for Fault Detection & Diagnostics Tools

Fault detection and diagnostics (FDD) analytical tools for heating, ventilation and air conditioning (HVAC) systems represent one of the most active areas of smart building technology development. A diversity of techniques is used for FDD analytics, spanning physical models, black box, and rule-based approaches, and researchers continuously strive to develop improved algorithms. With FDD algorithm numbers now in the hundreds, there is a need for performance evaluation of these algorithms in order to assess improvements, improve costeffectiveness, and to prioritize investment in the further development of these technologies. A persistent challenge of FDD advance has been the lack of common datasets to benchmark the performance accuracy of FDD algorithms. This paper summarizes the successful curation of HVAC operational data, paired with validated ground-truth information regarding the presence and absence of faults. The current dataset, consisting of both simulation and experimental data, will evolve to include a larger set of HVAC systems with the objective of creating the largest publicly available dataset to be used by FDD developers, users, and researchers to compare and contrast performance accuracy across FDD algorithms, helping to drive improvements that will spur greater market adoption of FDD tools. Furthermore, in order to avoid previously observed issues with contributed datasets and ensure high quality and consistency of future submissions, the development of data validation and ground-truth assessment protocol is detailed in this study.

Casillas, Armando↗

INTERFUEL: FAST - FY 2020 Federal Fleet Dataset [Slides]

This presentation provides an overview of the fiscal year (FY) 2020 federal motor vehicle fleet dataset collected through the Federal Automotive Statistical Tool (FAST). FAST is a web-based information system sponsored by GSA's Office of Government-wide Policy and DOE's Federal Energy Management Program to collect information about the US federal government's fleet of motor vehicles. The presentation discusses the size of the dataset; a high-level look at what the information collected shows about the makeup and operation of the federal motor vehicle fleet during FY 2020 and how that compares with recent years; the process used to review the agency submissions comprising the dataset and the impact of some of the agency-provided corrections to identified issues; activities following the finalization of the dataset.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Dataset Repository for Investigating Suicide Risk Using Social and Environmental Determinants of Health

Suicide is frequently modeled as a function of genetics and environment, where the latter refers to factors other than direct biological consequences, such as air quality, financial level, social connectivity, transportation and food access, and homelessness status. According to the World Health Organization, clean air, a stable climate, adequate water, sanitation and hygiene, safe chemical use, radiation protection, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved natural environment are all prerequisites for good health. Understanding the relationships between these determinants and mental health outcomes requires standardized data that can be included in healthcare programs and health outcome models. There is a wealth of publicly available data on social and environmental factors provided by various US organizations that can benefit the design of health care systems and public health interventions, as well as improve our comprehension of factors that impact health. Such information would not only help improve the understanding of individual and community risk but also identify new risk factors that have not previously been therapeutically targeted, especially in terms of their impact on mental health. However, curating and standardizing such datasets is challenging because they are often recorded at numerous geographical and temporal resolutions and with varying spatial and temporal granularities. To address this challenge, we launched an endeavor in conjunction with the Veterans Health Administration to collect publicly available socioeconomic and environmental determinants of health statistics in the US. In this manuscript, we describe a social and environmental determinants of health (SEDH) datasets repository, data curation documentation, and a pipeline framework for data generation; This effort started in 2020, when we began constructing a scalable pipeline to automate the download, extraction, preparation, analysis, and production of datasets. These datasets have been made available to the VHA and may be shared upon agreement with collaborating organizations.

60 APPLIED LIFE SCIENCES↗