Automated Classification and Validation System for Building Data.
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
In preparation for the second runs of the ProtoDUNE detectors at CERN (NP02 and NP04)[1], DUNE has established a new data pipeline for bringing the data from the EHN-1 experimental hall at CERN to primary tape storage at Fermilab and CERN, and then spreading it out to a distributed disk data store at many locations around the world. This system includes a new Ingest Daemon and a new Declaration Daemon. The Rucio[2] replica catalog, and FTS3 transport are used to transport all files. All file metadata is declared to the new MetaCat[3] metadata service. All of these new components have been successfully tested at a scale equal to the expected output of the detector data acquisition system (~2-4 GB/s), and the expected network bandwidth out of the experimental hall. We present the procedure that was used to test and the results of the test.
This paper introduces a database of 34 field-measured building occupant behavior datasets collected from 15 countries and 39 institutions across 10 climatic zones covering various building types in both commercial and residential sectors. This is a comprehensive global database about building occupant behavior. The database includes occupancy patterns (i.e., presence and people count) and occupant behaviors (i.e., interactions with devices, equipment, and technical systems in buildings). Brick schema models were developed to represent sensor and room metadata information. The database is publicly available, and a website was created for the public to access, query, and download specific datasets or the whole database interactively. The database can help to advance the knowledge and understanding of realistic occupancy patterns and human-building interactions with building systems (e.g., light switching, set-point changes on thermostats, fans on/off, etc.) and envelopes (e.g., window opening/closing). With these more realistic inputs of occupants’ schedules and their interactions with buildings and systems, building designers, energy modelers, and consultants can improve the accuracy of building energy simulation and building load forecasting.
Deep learning models have shown promise in reservoir inflow prediction, yet their performance often deteriorates when applied to different reservoirs due to distributional differences, referred to as the domain shift problem. Domain generalization (DG) solutions aim to address this issue by extracting domain-invariant representations that mitigate errors in unseen domains. However, in hydrological settings, each reservoir exhibits unique inflow patterns, while some metadata beyond observations like spatial information exerts indirect but significant influence. This mismatch limits the applicability of conventional DG techniques to many-domain hydrological systems. To overcome these challenges, we propose HydroDCM, a scalable DG framework for cross-reservoir inflow forecasting. Spatial metadata of reservoirs is used to construct pseudo-domain labels that guide adversarial learning of invariant temporal features. During inference, HydroDCM adapts these features through light-weight conditioning layers informed by the target reservoir’s metadata, reconciling DG’s invariance with location-specific adaptation. Experiment results on 30 real-world reservoirs in the Upper Colorado River Basin demonstrate that our method substantially outperforms state-of-the-art DG baselines under many-domain conditions and remains computationally efficient.
This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.
As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.
This paper explores the challenges and solutions for managing and processing the vast amount of data generated by the Advanced Photon Source (APS), a synchrotron light source facility producing ultra-bright x-rays for diverse scientific domains. With 68 experimental beamlines covering materials research, biology, and more, the APS serves a wide user base across academia, government, and industry. The ongoing upgrade of the APS storage ring and installation of new instruments will amplify data generation and processing demands. This paper discusses the approach to address these demands through automated data processing using standardized workflows that produce faster scientific insights. The APS Data Management System coordinates various data related tasks to manage storage, data transfer, metadata cataloging, data processing, and interfaces with tools provided by Globus. Through integration with the Argonne Leadership Computing Facility (ALCF), APS users can efficiently access high-performance computing resources. Standardized workflows have led to reduced computational burdens on scientists and greater accessibility of high performance computing resources. We demonstrate how standardization and collaboration enable scientists to rapidly convert raw data into meaningful scientific results, establishing a streamlined path from data collection to analysis and ultimately to publication.
This dataset provides Level 1 (L1) discrete-return light detection and ranging (LiDAR) point cloud data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary unclassified discrete-return LiDAR data delivered by NEON and are provided per flightline as LASzip (LAZ) 1.4 Format 6 files. Data were processed following the workflow described in the NEON L0-to-L1 Discrete Return LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022). Each record in the unclassified point clouds represents a geolocated laser target/return recorded by the LiDAR system, with values for X, Y, Z position and return intensity. All point coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Flight metadata describing flightline boundaries and positional uncertainty by point are also included. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.
Surface soil moisture (mrsos) and vertically integrated soil moisture (mrsol) over the top 10 cm should, by definition, be physically consistent in Earth System Models (ESMs). However, an evaluation of nine CMIP6 models reveals substantial inconsistencies: in some models, mrsos and integrated mrsol agree globally; in others, they align only in specific regions; and in a few, they diverge across all grid cells. These discrepancies arise from a combination of factors, including metadata errors, inconsistent variable definitions, or diagnostic sequencing within the model. We demonstrate how such issues can lead to significant biases, even when both variables are present and seemingly well-defined. As model complexity increases and multi-model comparisons become more common, assumptions about variable equivalence may lead to flawed conclusions. This study highlights the need for routine consistency checks, improved metadata standards, and community-wide practices that ensure reliability of derived variables across ESM outputs, particularly in preparation for CMIP7.
Machine Learning (ML) has become a critical tool enabling new methods of analysis and driving deeper understanding of phenomena across scientific disciplines. There is a growing need for "learning systems" to support various phases in the ML lifecycle. While others have focused on supporting model development, training, and inference, few have focused on the unique challenges inherent in science, such as the need to publish and share models and to serve them on a range of available computing resources. In this paper, we present the Data and Learning Hub for science (DLHub), a learning system designed to support these use cases. Specifically, DLHub enables publication of models, with descriptive metadata, persistent identifiers, and flexible access control. It packages arbitrary models into portable servable containers, and enables low-latency, distributed serving of these models on heterogeneous compute resources. In this work, we show that DLHub supports low-latency model inference comparable to other model serving systems including TensorFlow Serving, SageMaker, and Clipper, and improved performance, by up to 95%, with batching and memoization enabled. We also show that DLHub can scale to concurrently serve models on 500 containers. Finally, we describe five case studies that highlight the use of DLHub for scientific applications.
Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields
Not Available
Infrastore is time-series storage for energy-systems simulations, backed by HDF5 + SQLite, with Rust, Python, Julia, gRPC, and CLI bindings. It is a Rust library for managing time-series data in power-systems and energy simulations. Numerical arrays are persisted in HDF5, and the metadata associating each array with its owning component lives in SQLite. Identical arrays are stored once and shared through content addressing. It ships native Rust, Python (PyO3), and Julia (C ABI) interfaces, the infrastore command-line tool, and a read-only gRPC server with a Rust client. Documentation: https://natlabrockies.github.io/infrastore/latest/ — start with the Quick Start or the Architecture.
The following information and metadata applies to both the Phase I (Hydrodynamics) and Phase II (Full System Power Take-Off) zip folders which contain testing data from the OSU (Oregon State University) O.H. Hinsdale Wave Research Laboratory, from both OSU and the University of Hawaii at Manoa (UH). See zip folders provided further below in the downloads section. For experimental data of the full system, including PTO, see Phase II dataset. There are two main directories in each Phases's zip folder: "OSU_data" and "UH_data". The "OSU_data" directory contains data collected from their DAQ (data acquisition system), which includes all wave gauge observations, as well as body motions derived from their Qualisys motion tracking system. The organization of the directory follows OSU's convention. Detailed information on the instrument setup can be found under "OSU_data/docs/setup/instm_locations". The experiments conducted are documented in the "OSU_data/docs/daq_logs", which provides the trial number to the corresponding data located under "OSU_data/data" in several formats (e.g., ".mat" and ".txt"). Inside the trial directory, data is provided for each of the instruments defined in "OSU_data/docs/setup/instm_locations". The "UH_data" directory contains data collected from their DAQ. The data is stored in a ".tdms" file format. There are free plug-ins for Microsoft Excel and MathWorks MATLAB to read the ".tdms" format. Below are a few links providing methods to read in the data, but a Google search should identify alternatives sources if these no longer exist (valid as of January 2024): Excel: http://www.ni.com/example/27944/en/ MATLAB: https://www.mathworks.com/matlabcentral/fileexchange/30023-tdms-reader The Excel plugin is recommend to get a quick overview of the data. The UH data is organized by directory name, in which the sub-directories for each experiment contains a directory whose name defines the wave height and period for the experimental data within. For example, a directory name "H02_T0275" corresponds to an experiment with wave height 0.1m and a period of 2.75s. For random wave data, the gamma value is also included in the directory name. For example, a directory name "H02_T0225_G18" corresponds to an experiment with a significant wave height of 0.2m, a peak period of 2.25s, and a gamma value of 1.8, with each spectra being a TMA spectrum. For the free decay experiments, the directory name is defined by the initial angular displacement. For example, a directory name "ang05_run01" corresponds to an experiment with an initial angular displacement of 5 degrees. There is a dataset in the UH data for each corresponding experiment defined in the OSU DAQ logs. The ".tdms" data is output from the DAQ at fixed intervals. Therefore, if multiple files are contained within the folder, the data will need to be stitched together. Within the UH dataset, there are two input channels from the OSU DAQ providing a random square wave signal for time synchronization ("ENV-WHT-0010") and a high/low signal ("ENV-WHT-0012") to identify when the wave maker is active (+5V). The UH data is logged as a collection of channel outputs. Channels not in use for the OSU testing (either Phase I or Phase II) are marked "nan" below. If the sensor is disconnected, it will record noise throughout the experiment. Below are the channel definitions in terms of what they measure: GPS Time = time CYL-POS-0001 = position between flap and fixed reference CYL-LCA-0001 = force between flap and hydraulic cylinder REC-LPT-0001 = nan REC-HPT-0001 = nan REC-HPT-0002 = nan REC-HPT-0003 = nan HHT-HPT-0001 = pressure at exhaust ("head" only) REC-FQC-0001 = nan REC-FQC-0002 = nan HHT-FQC-0001 = flow at exhaust ("head" only) ENV-WHT-0001 = nan ENV-WHT-0002 = nan ENV-WHT-0003 = nan ENV-WHT-0010 = random signal from OSU DAQ ENV-WHT-0012 = high/low signal from OSU DAQ Also included is a calibration curve to convert the string pot data to flap pi...
Organizations that monitor for underground nuclear explosive tests are interested in techniques that automatically characterize mining blasts to reduce the human analyst effort required to produce high - quality event bulletins. Waveform correlation is effective in finding similar waveforms from repeating seismic events, including mining blasts. In this study we use waveform template event metadata to seek corroborating detections from multiple stations in the International Monitoring System of the Preparatory Commission for the Comprehensive Nuclear-Test-Ban Treaty Organization. We build upon events detected in a prior waveform correlation study of mining blasts in two geographic regions, Wyoming and Scandinavia. Using a set of expert analyst-reviewed waveform correlation events that were declared to be true positive detections, we explore criteria for choosing the waveform correlation detections that are most likely to lead to bulletin-worthy events and reduction of analyst effort.
Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).
Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).
Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.