Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “building data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

What We Learned From Analyzing 18 Million Rows of Commercial Buildings’ HVAC Fault Data

To achieve ambitious decarbonization goals it is critical that buildings operate to their full potential. Commercial HVAC systems, however, experience a wide range of operational faults, adversely affecting energy consumption, occupant comfort, and maintenance costs. Analytical tools such as fault detection & diagnostics (FDD) software identify and help diagnose these types of sensing, mechanical, or control-related faults. While significant energy savings has been documented for FDD, along with limited-scale studies on technical capabilities, there is a lack of empirical data on faults being reported by FDD tools. With FDD deployment accelerating significantly over the past decade there is an opportunity to gather and analyze data on commercial HVAC operational problems at an unprecedented scale. Such data could address many questions such as: [a] What faults are most commonly reported?; and [b] How does fault reporting vary by time of year and other possible drivers? A recent study into FDD fault reporting amassed the largest U.S. dataset of commercial HVAC air-side fault records, drawn from multi-year monitoring across over 60,000 pieces of HVAC equipment. The results of this study provide granular data on fault reporting for over 90 unique fault types. In this paper we provide an overview of the research process and highlight key findings and lessons learned. This study presents an extraordinary level of detail on FDD fault reporting characteristics across many climate zones and building types. Armed with these new insights, commercial building industry stakeholders can make better informed decisions when designing, configuring, and operating commercial HVAC systems.

Crowe, Eliot↗

Data-Driven Modeling and Optimization of Building Energy Consumption: a Case Study

Installing sensors and Building Automation Systems (BAS) allows controlling the facility operations while generating data that can be analyzed for model development. This work focuses on data-driven modeling of the building to optimize energy consumption. The City of Orlando aims to reduce its energy consumption so, they provided us access to their BAS for data and studying the operation of its facilities. We selected a mid-size pilot building to conduct data analysis and modeling. We develop an Application Programming Interface (API) to login to the servers and scrape data. The scraped data contains features ranging from environmental conditions to equipment activity. This dataset is a time series so, it's handled in accordance and analyzed to investigate patterns and relations between data points that help choose parameters for predictive models for building and equipment. Finally, the models are optimized to reduce the energy consumption of the facility.

Grover, Divas↗

Using Open Data to Characterize Building Stock Trends for Energy and Equity Evaluation

This project provides a methodology for retrieving data from both existing public data and municipal open data sets to characterize changes in the building stock, socioeconomic, demographic, and climate indicators at a neighborhood scale. This methodology forms the basis for a geospatial dashboard to analyze and share the data. The developed methodology may be used to collect data for further energy equity analyses. Additional data sets, such as from other municipalities, may be incorporated in the future to expand the extents of the possible analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Utility-scale Building Type Assignment Using Smart Meter Data

United States building energy use accounted for 40% of total energy use, 74% of peak demand, and $412 billion in 2019. Building energy modeling allows researchers to simulate building physics, gain insights into possible energy/demand saving opportunities, and assess cost-effective resilience amidst climate change. Many building features needed to create building energy models are readily available such as 2D footprints and LiDAR (height). A critical feature that is not generally obtainable is the building type. In partnership with a utility, a years worth of real-world, 15-minute electrical use data has been examined. The smart meter data is compared to 97 different prototype building energy models to assign building type. Real-world considerations including data preparation, quality assurance, and handling of missing values for advanced metering infrastructure data are addressed. Euclidean distance for pattern-matching of energy use, dynamic time warping, and time-window statistics with machine learning are compared for determining building type from measured electricity use.

Bass, Brett↗

Automatic Segmentation of Building Envelope Point Cloud Data Using Machine Learning

About 50% of buildings in the US were constructed before energy codes were introduced. Modular overclad panel retrofits, in which a new envelope is constructed over the existing building, are a promising solution given that it minimizes occupant disruption and shortens construction time at the jobsite. Current state-of-the-art retrofit panel layout and dimensioning consists of three steps: 1) 3D point cloud data generation of the building envelope using commonly available surveying equipment, 2) manual segmentation of 3D point cloud data by a trained professional to identify and dimension window openings, door openings, and other architectural features, and 3) modular panel layout optimization and dimensioning by an architect or engineer. Among these steps, the second one remains the most difficult and costly because it is very labor-intensive. We propose a methodology to automatically label 3D point cloud data to reduce the time and expense spent in manual segmentation. Machine learning methods were employed to classify the point cloud data into distinct groups, each of which corresponds to different features of the building envelope. After classification, a segmentation algorithm was developed to perform boundary detection and separate the components of the façade. Finally, the algorithm returns the relative positions and dimensions of the features in the building envelope. The measurements obtained with the proposed automated method were compared against the actual dimensions to determine the overall algorithm accuracy. The proposed algorithm can then be used to reduce manual efforts for 3D point cloud labeling before modular panel layout optimization is performed.

Maldonado Puente, Bryan↗

Chicago microclimate and building energy use data

These data comprise three elements: - High resolution, 90 m simulated weather data for 1 year at 15 min. intervals (with known gaps toward the end of each month). These files are in .csv format. - A mapping of individual buildings with individual IDs, their latitude/longitude location, and height. (Excel file) - Energy simulation output of these individual buildings, at 15 min. intervals for a whole year. (.json and other files)

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Chicago microclimate and building energy use data

These data comprise three elements: - High resolution, 90 m simulated weather data for 1 year at 15 min. intervals (with known gaps toward the end of each month). These files are in .csv format. - A mapping of individual buildings with individual IDs, their latitude/longitude location, and height. (Excel file) - Energy simulation output of these individual buildings, at 15 min. intervals for a whole year. (.json and other files)

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Processed data from the Building Management System for the System Engineering Building.

The dataset spans November 2018 to May 2020 and includes time-series measurements corresponding to supply and return temperatures of air and water, air, hot water and cold water flow rates, energy and power consumption, set-points etc. as a single CSV file. In addition to the measurements, a metadata .json file, and a .ttl file to visualize the data as per BRICK schema are also included.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Unified architecture for data-driven metadata tagging of building automation systems

This article presents a Unified Architecture (UA) for automated point tagging of Building Automation System (BAS) data, based on a combination of data-driven approaches. Advanced energy analytics applications—including fault detection and diagnostics and supervisory control—have emerged as a significant opportunity for improving the performance of our built environment. Effective application of these analytics depends on harnessing structured data from the various building control and monitoring systems, but typical BAS implementations do not employ any standardized metadata schema. While standards such as Project Haystack and Brick Schema have been developed to address this issue, the process of structuring the data, i.e., tagging the points to apply a standard metadata schema, has, to date, been a manual process. This process is typically costly, labor-intensive, and error-prone. In this work we address this gap by proposing a UA that automates the process of point tagging by leveraging the data accessible through connection to the BAS, including time-series data and the raw point names. The UA intertwines supervised classification and unsupervised clustering techniques from machine learning and leverages both their deterministic and probabilistic outputs to inform the point tagging process. Furthermore, we extend the UA to embed additional input and output data-processing modules that are designed to address the challenges associated with the real-time deployment of this automation solution. We test the UA on two datasets for real-life buildings: (i) commercial retail buildings and (ii) office buildings from the National Renewable Energy Laboratory (NREL) campus. We report the proposed methodology correctly applied 85–90% and 70–75% of the tags in each of these test scenarios, respectively for two significantly different building types used for testing UA's fully-functional prototype. The proposed UA, therefore, offers promising approach for automatically tagging BAS data as it reaches close to 90% accuracy. Further building upon this framework to algorithmically identify the equipment type and their relationships is an apt future research direction to pursue.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Prevalence of typical operational problems and energy savings opportunities in U.S. commercial buildings

In the United States, as much as 30% of the 19 EJ that commercial buildings consume is considered excess. Much of the excess energy is due to the inability to manage building operations efficiently. Because almost 20% of the total primary energy consumption is associated with commercial buildings, significant energy reductions in this sector are needed to mitigate climate change. Therefore, many cities and states are mandating periodic “tune-ups” of these buildings to eliminate excess energy consumption. Although the benefits of tune-ups and retro-commissioning are clear, focusing these mandates to look for specific opportunities has been a challenge because of the lack of studies that document the prevalence of opportunities. Therefore, we analyzed building automation system data from 151 buildings across the United States to document common operational problems and opportunities to improve building operations. This analysis showed that opportunities to improve building operations exist in almost every building. These opportunities were not strongly correlated with building vintage or size, but were reflective of how the buildings are operated. The prevalence of the top 20 opportunities ranged between 74% and 23%, with 40% of these associated with air-handling units. The rest of the opportunities are associated with schedules, chilled and hot-water distribution, and zone controls. Of the 151 buildings, 69 of them implemented corrective actions of some or all opportunities that were identified. Implementation varied across the Re-tuning categories, with 60% for schedule opportunities, 50% for zone opportunities, over 40% for the air-handling unit and hot-water opportunities, and 35% of the chilled-water opportunities. There was wide variation in whole building energy savings, ranging from 0 to 50% and 0 to 18 $/m2 with median percent annual whole building savings of 12% and median normalized annual cost savings of $1.75/m2. In addition to documenting these key findings, the paper provides a list of opportunities that can be automatically and continuously identified and corrected and offers a list of those opportunities that should be the focus of the mandates.

Katipamula, Srinivas↗

Selective Sampling for Sensor Type Classification in Buildings

A key barrier to applying any smart technology to a building is the requirement of locating and connecting to the necessary resources among the thousands of sensing and control points, i.e., the metadata mapping problem. Existing solutions depend on exhaustive manual annotation of sensor metadata --- a laborious, costly, and hardly scalable process. To reduce the amount of manual effort required, this paper presents a multi-oracle selective sampling framework to leverage noisy labels from information sources with unknown reliability such as existing buildings, which we refer to as weak oracles, for metadata mapping. This framework involves an interactive process, where a small set of sensor instances are progressively selected and labeled for it to learn how to aggregate the noisy labels as well as to predict sensor types. Two key challenges arise in designing the framework, namely, weak oracle reliability estimation and instance selection for querying. To address the first challenge, we develop a clustering-based approach for weak oracle reliability estimation to capitalize on the observation that weak oracles perform differently in different groups of instances. For the second challenge, we propose a disagreement-based query selection strategy to combine the potential effect of a labeled instance on both reducing classifier uncertainty and improving the quality of label aggregation. We evaluate our solution on a large collection of real-world building sensor data from 5 buildings with more than 11,000 sensors of 18 different types. The experiment results validate the effectiveness of our solution, which outperforms a set of state-of-the-art baselines.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

TListSpectrum

TListSpectrum is a C++ class developed inside the CERN high-energy physics analysis C++ framework ROOT. This class structure was developed to assist in the processing, visualization, and analysis of list-mode or time-stamped radiation spectroscopy data. The class structure currently contains parsing and functionality to synthesize list-mode data from CAEN and Mirion Lynx radiation spectroscopy digital acquisition systems along with feature functionality to post-process data sets and build coincident data sets from the instrument.

Pierson, Bruce↗

Model America - data and models of every U.S. building

The 5-year goal of the 'Model America' concept was to generate a model of every building in the United States. This data repository delivers on that goal. Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,714,640 buildings detected in the United States and this dataset contains 122,930,327 (97.8%) buildings which resulted in a successful simulation. Future, annual updates have been proposed that may include additional buildings, data improvements, or other algorithmic enhancements. This dataset of 122.9 million buildings includes: Models (state_county.zip) - OpenStudio (v3.1.0) and EnergyPlus (v9.4) building energy models. Please note that the download requires the free Globus Connect Personal (https://www.globus.org/globus-connect-personal); Each model has approximately 3,000 building input descriptors that can be extracted. Please see the EnergyPlus(v9.4) 2,784-page Input/Output Reference Guide (https://energyplus.net/sites/all/modules/custom/nrel_custom/pdfs/pdfs_v9.4.0/InputOutputReference.pdf) for everything that can be retrieved or simulated from these models. These models were derived from the following metadata, which is not included in this dataset: 1. ID - unique building ID 2. County - county name 3. State - state name 4. CZ - ASHRAE Climate Zone designation 5. Clim_Zone - text label of climate zone 6. est_year - estimated year of construction 7. est_commercial - estimated building type (0=residential, 1=commercial) 8. Centroid - building center location in latitude/longitude (from Footprint2D) 9. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 10. Height - building height (meters) 11. Area2D - footprint area (ft2) 12. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 13. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 14. NumFloors - number of floors (above-grade) 15. Area - estimate of total conditioned floor area (ft2) 16. Standard - building vintage. These models are made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy's (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). This research used resources of the Argonne Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC02-06CH11357. Please cite as: New, Joshua R., Adams, Mark, Bass, Brett, Berres, Anne, and Clinton, Nicholas (2021). 'Model America - data and models of every U.S. building. [Data set].' Constellation, doi.ccs.ornl.gov/ui/doi/339, April 14, 2021

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unique Building Identifier (UBID): Public Sector Implementation Guide

Buildings generate data throughout their lifecycle – about ownership & taxation, usage, zoning, code compliance, energy use, and retrofits. State and local governments collect this data after it flows through growing networks of people and systems. But collecting data is only half the battle; what’s really needed is information – the actionable insights that lead to successful policy outcomes.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Model America: Data and Models for every U.S. Building

The 5-year goal of the “Model America” concept was to generate a model of every building in the United States. This data repository delivers on that goal with "Model America v1". Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,715,609 buildings detected in the United States. Of this number, 122,146,671 (97.2%) buildings resulted in a successful generation and simulation of a building energy model. This dataset includes the full 125 million buildings. Future updates may include additional buildings, data improvements, or other algorithmic model enhancements in "Model America v2". This dataset contains OSM and IDF zip files for every U.S. county. Each zip file contains the generated buildings from that county. The .csv input data contains the following data fields: 1. ID - the Unique Building Identifier (UBID), generated using the Pacific Northwest National Laboratory (PNNL) BuildingID framework 2. Centroid - building center location in latitude/longitude (from Footprint2D) 3. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 4. State_abbr - state name 5. Area - estimate of total conditioned floor area (ft2) 6. Area2D - footprint area (ft2) 7. Height - building height (ft) 8. NumFloors - number of floors (above-grade) 9. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 10. CZ - ASHRAE Climate Zone designation 11. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 12. Standard - building vintage This data is made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy’s (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). Update (September 23, 2025): We corrected the ID field in all state-level.csv input files to ensure one-to-one consistency with the corresponding .osm and .idf output files. The schema and file structure are unchanged; only the values in the ID column were modified. No files were added or removed, and the .zip bundles (containing .osm / .idf) are unchanged. The corrected .csv inputs were re-extracted in March 2025 from the original data generated ~ 2021 (Theta supercomputer runs), and published here to align input IDs with model outputs. Update (September 6, 2026): The Model America dataset was updated to replace the previous building ID field with the Unique Building Identifier (UBID), using the Pacific Northwest National Laboratory (PNNL) BuildingID framework. UBIDs provide standardized, location-based identifiers for individual building footprints and improve interoperability with other building and geospatial datasets. The data files containing the previous building identifiers were updated to include UBIDs. This update standardizes building identification; the underlying Model America building characteristics and energy simulation results were not recomputed as part of this update.

54 ENVIRONMENTAL SCIENCES↗