Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “building data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

TListSpectrum

TListSpectrum is a C++ class developed inside the CERN high-energy physics analysis C++ framework ROOT. This class structure was developed to assist in the processing, visualization, and analysis of list-mode or time-stamped radiation spectroscopy data. The class structure currently contains parsing and functionality to synthesize list-mode data from CAEN and Mirion Lynx radiation spectroscopy digital acquisition systems along with feature functionality to post-process data sets and build coincident data sets from the instrument.

Pierson, Bruce↗

Model America - data and models of every U.S. building

The 5-year goal of the 'Model America' concept was to generate a model of every building in the United States. This data repository delivers on that goal. Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,714,640 buildings detected in the United States and this dataset contains 122,930,327 (97.8%) buildings which resulted in a successful simulation. Future, annual updates have been proposed that may include additional buildings, data improvements, or other algorithmic enhancements. This dataset of 122.9 million buildings includes: Models (state_county.zip) - OpenStudio (v3.1.0) and EnergyPlus (v9.4) building energy models. Please note that the download requires the free Globus Connect Personal (https://www.globus.org/globus-connect-personal); Each model has approximately 3,000 building input descriptors that can be extracted. Please see the EnergyPlus(v9.4) 2,784-page Input/Output Reference Guide (https://energyplus.net/sites/all/modules/custom/nrel_custom/pdfs/pdfs_v9.4.0/InputOutputReference.pdf) for everything that can be retrieved or simulated from these models. These models were derived from the following metadata, which is not included in this dataset: 1. ID - unique building ID 2. County - county name 3. State - state name 4. CZ - ASHRAE Climate Zone designation 5. Clim_Zone - text label of climate zone 6. est_year - estimated year of construction 7. est_commercial - estimated building type (0=residential, 1=commercial) 8. Centroid - building center location in latitude/longitude (from Footprint2D) 9. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 10. Height - building height (meters) 11. Area2D - footprint area (ft2) 12. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 13. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 14. NumFloors - number of floors (above-grade) 15. Area - estimate of total conditioned floor area (ft2) 16. Standard - building vintage. These models are made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy's (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). This research used resources of the Argonne Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC02-06CH11357. Please cite as: New, Joshua R., Adams, Mark, Bass, Brett, Berres, Anne, and Clinton, Nicholas (2021). 'Model America - data and models of every U.S. building. [Data set].' Constellation, doi.ccs.ornl.gov/ui/doi/339, April 14, 2021

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unique Building Identifier (UBID): Public Sector Implementation Guide

Buildings generate data throughout their lifecycle – about ownership & taxation, usage, zoning, code compliance, energy use, and retrofits. State and local governments collect this data after it flows through growing networks of people and systems. But collecting data is only half the battle; what’s really needed is information – the actionable insights that lead to successful policy outcomes.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Model America: Data and Models for every U.S. Building

The 5-year goal of the “Model America” concept was to generate a model of every building in the United States. This data repository delivers on that goal with "Model America v1". Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,715,609 buildings detected in the United States. Of this number, 122,146,671 (97.2%) buildings resulted in a successful generation and simulation of a building energy model. This dataset includes the full 125 million buildings. Future updates may include additional buildings, data improvements, or other algorithmic model enhancements in "Model America v2". This dataset contains OSM and IDF zip files for every U.S. county. Each zip file contains the generated buildings from that county. The .csv input data contains the following data fields: 1. ID - the Unique Building Identifier (UBID), generated using the Pacific Northwest National Laboratory (PNNL) BuildingID framework 2. Centroid - building center location in latitude/longitude (from Footprint2D) 3. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 4. State_abbr - state name 5. Area - estimate of total conditioned floor area (ft2) 6. Area2D - footprint area (ft2) 7. Height - building height (ft) 8. NumFloors - number of floors (above-grade) 9. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 10. CZ - ASHRAE Climate Zone designation 11. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 12. Standard - building vintage This data is made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy’s (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). Update (September 23, 2025): We corrected the ID field in all state-level.csv input files to ensure one-to-one consistency with the corresponding .osm and .idf output files. The schema and file structure are unchanged; only the values in the ID column were modified. No files were added or removed, and the .zip bundles (containing .osm / .idf) are unchanged. The corrected .csv inputs were re-extracted in March 2025 from the original data generated ~ 2021 (Theta supercomputer runs), and published here to align input IDs with model outputs. Update (September 6, 2026): The Model America dataset was updated to replace the previous building ID field with the Unique Building Identifier (UBID), using the Pacific Northwest National Laboratory (PNNL) BuildingID framework. UBIDs provide standardized, location-based identifiers for individual building footprints and improve interoperability with other building and geospatial datasets. The data files containing the previous building identifiers were updated to include UBIDs. This update standardizes building identification; the underlying Model America building characteristics and energy simulation results were not recomputed as part of this update.

54 ENVIRONMENTAL SCIENCES↗

Educational Consortium for Energy-related Data Science & Computation in Building Engineering Programs

The project spearheaded by Pennsylvania State University aims to address the growing need for integrating energy-focused computation and data science into building engineering education. As the demand for energy-efficient building designs and operations increases, the educational sector must adapt to equip future engineers with the necessary skills. This initiative responds to this need by developing a consortium that unites multiple institutions to enhance curriculum development, dataset curation, and resource sharing, thereby ensuring students are well-prepared for the evolving energy sector. The primary goal of the project is to establish a consortium that will develop and disseminate educational materials and training programs focused on energy-related data science and computation. Key accomplishments include the creation of a beta website for resource sharing, the development of training programs and standalone modules, and the curation of datasets accessible to the public. This effort will culminate in a curriculum that incorporates advanced modeling technologies and data science skills into building engineering programs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Calibration of urban building energy model using smart meter data for district peak load prediction

Urban building energy modeling (UBEM) is a powerful approach to assessing baseline building energy performance and retrofits with new technologies across building stocks in cities. However, the accuracy of UBEM is often constrained by the limited availability of reliable data about building characteristics and operations, such as envelope efficiency levels, HVAC system performance, and end-use load patterns. Existing research has performed UBEM calibration using annual or monthly energy consumption data, which falls short when higher-resolution time series applications are needed, such as peak load prediction for utility operation planning. This study presents a new framework for calibrating building energy models at urban scale using smart meter data, targeting the accurate prediction of summer peak electricity loads to support robust grid planning. The framework first integrates various data sources to enhance baseline input assumptions for building models, and then calibrates the baseline models through a pattern-matching approach. A case study using CityBES and two years of AMI data from over 9000 residential customers in Portland, Oregon, demonstrated the workflow and its effectiveness. The calibrated models achieved a daily peak load mean absolute percentage error of 2.6 % during the heatwave in the calibration year, and 2.0 % in the validation year using another year of AMI data. Using the calibrated models, we analyzed the demand flexibility potential of the district building stock as an application of UBEM calibration. The findings affirm the appropriate use of UBEM for peak electric load forecasting and demand side management at the utility distribution system level.

AMI data↗

Building Analytics Tool Deployment at Scale: Benefits, Costs, and Deployment Practices

Buildings are becoming more data-rich. Building analytics tools, including energy information systems (EIS) and fault detection and diagnostic (FDD) tools, have emerged to enable building operators to translate large amounts of time-series data into actionable findings to achieve energy and non-energy benefits. To expedite data analytics adoption and facilitate technology innovation, building owners, technology developers, and researchers need reliable cost–benefit data and evidence-based guidance on deployment practices. This paper fulfills these needs with the energy use and survey data from a wide-ranging research and industry partnership program that covers thousands of buildings installed with analytics tools. The paper indicates that after two years of implementation, organizations using FDD tools and EIS tools achieved 9% and 3% median annual energy savings, respectively. The median base cost and annual recurring cost for FDD are USD 0.65 per square meter (m2) (USD 0.06 per square foot [ft2]) and USD 0.22 per m2 (USD 0.02 per ft2), and are USD 0.11 per m2 (USD 0.01 per ft2) and USD 0.11 per m2 (USD 0.01 per ft2) for EIS. The common metrics and analyses that are used in the tools to support the discovery of energy efficiency measures are summarized in detail. Two best practice examples identified to maximize the benefits of tool implementation are also presented. Opportunities to advance the state of technology include simplified data integration and management, and more efficient processes for acting on analytics outputs. Compared with previous efforts in the literature, the findings presented in this paper demonstrate the effectiveness of building analytics tools with the largest known dataset.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Artificial Intelligence for Enhancing Multiscale Analysis: Buildings Focus

This project aims to develop multi-scale building energy data, potentially improving the representation of the U.S. buildings sector in GCAM-USA, an U.S.-focused human-energy-Earth systems model. Existing building energy datasets are typically limited to national or regional levels, which constrains the ability of models to capture fine-scale human-energy-Earth systems interactions and reduces their relevance for decision-making on issues such as energy security, resilience, and energy planning. By leveraging AI and advanced data integration methods, this work fuses multiple existing datasets to enhance the physical and geographic representation of both residential and commercial building energy use. So far, progress includes processing residential building data, designing the data structure for commercial buildings, and testing AI approaches for integrating datasets and addressing spatial-temporal gaps. This effort can not only advances GCAM-USA’s capability in modeling the buildings sector but also supports broader DOE missions, such as developing digital testbeds, enhancing grid resilience analysis, and improving building–energy system modeling at decision-relevant scales.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Machine Learning for Automated Extraction of Building Geometry

As data science comes to buildings, the promise of using machine learning and novel sources of data has received much attention. Advances in machine learning and computer vision algorithms, combined with increased access to unstructured data (e.g., images and text), have created an opportunity for automated extraction of building characteristics – cost-effectively, and at scale. Acquisition of features such as footprint are time consuming and costly to acquire with today’s manual methods, but can be streamlined through intelligent software-based solutions applied to satellite images. When combined with aerial RGB and thermal images, full 3D geometries and thermal maps can be constructed to determine additional characteristics such as window to wall ratio, height, number of stories and envelope thermal characteristics. In this paper we present three contributions to accelerate these high potential opportunities: (1) a methodical analysis of how these features can be integrated into today’s simulation and data driven software tools to enhance efficiency measure identification and owner/operator decision making; (2) development and accuracy testing of open source deep neural network methods to extract building footprints from satellite imagery, including the curation and application of openly available GIS datasets for training and continued development by others; and (3) an open framework for drone-based image capture and creation of 3D building geometries. This work represents an important bridge between high-level studies that span diverse application areas and those that detail point solutions yet cannot be easily replicated or extended.

Touzani, Samir↗

PV Window Building Study

Data associated with the paper “Photovoltaic Windows to Offset the Intensive Energy and Carbon Footprints of Highly Glazed Buildings” by Vincent M. Wheeler, Janghyun Kim, Tom Daligault, Bryan Rosales, Chaiwat Engtrakul, Robert C. Tenent, and Lance M. Wheeler

14 SOLAR ENERGY↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Data items will occasionally be removed from OSM if they are misidentified, if they no longer exist, if they are duplicates of another item, or similar. For that reason, updated versions of this database may not contain all data center locations included in previous versions. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Veridical data science

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, composed of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. As part of the PCS workflow, we develop PCS inference procedures, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others. Moreover, we demonstrate its favorable performance over existing methods in terms of receiver operating characteristic (ROC) curves in high-dimensional, sparse linear model simulations, including a wide range of misspecified models. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

97 MATHEMATICS AND COMPUTING↗

City and County Commercial Building Inventories

The Commercial Building Inventories provide modeled data on commercial building type, vintage, and area for each U.S. city and county. Please note this data is modeled and more precise data may be available through county assessors or other sources. Commercial building stock data is estimated using CoStar Realty Information, Inc. building stock data. This data is part of a suite of state and local energy profile data available at the "State and Local Energy Profile Data Suite" link below and builds on Cities-LEAP energy modeling, available at the "EERE Cities-LEAP Page" link below. Examples of how to use the data to inform energy planning can be found at the "Example Uses" link below.

Array↗

From RNNs to Foundation Models: An Empirical Study on Commercial Building Energy Consumption

Accurate short-term energy consumption forecasting for commercial buildings is crucial for smart grid operations. While smart meters and deep learning models enable forecasting using past data from multiple buildings, data heterogeneity from diverse buildings can reduce model performance. The impact of increasing dataset heterogeneity in time series forecasting, while keeping size and model constant, is understudied. We tackle this issue using the ComStock dataset, which provides synthetic energy consumption data for U.S. commercial buildings. Two curated subsets, identical in size and region but differing in building type diversity, are used to assess the performance of various time series forecasting models, including finetuned open-source foundation models (FMs). The results show that dataset heterogeneity and model architecture have a greater impact on post-training forecasting performance than the parameter count. Moreover, despite the higher computational cost, finetuned FMs demonstrate competitive performance compared to base models trained from scratch.

commercial buildings↗

Inferring building height from footprint morphology data

As cities continue to grow globally, characterizing the built environment is essential to understanding human populations, projecting energy usage, monitoring urban heat island impacts, preventing environmental degradation, and planning for urban development. Buildings are a key component of the built environment and there is currently a lack of data on building height at the global level. Current methodologies for developing building height models that utilize remote sensing are limited in scale due to the high cost of data acquisition. Other approaches that leverage 2D features are restricted based on the volume of ancillary data necessary to infer height. Here, we find, through a series of experiments covering 74.55 million buildings from the United States, France, and Germany, it is possible, with 95% accuracy, to infer building height within 3 m of the true height using footprint morphology data. Our results show that leveraging individual building footprints can lead to accurate building height predictions while not requiring ancillary data, thus making this method applicable wherever building footprints are available. The finding that it is possible to infer building height from footprint data alone provides researchers a new method to leverage in relation to various applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Aggregation and data driven identification of building thermal dynamic model and unmeasured disturbance

An aggregate model is a single-zone equivalent of a multi-zone building, and is useful for many purposes, including model based control of large heating, ventilation and air conditioning (HVAC) equipment. This paper deals with the problem of simultaneously identifying an aggregate thermal dynamic model and unknown disturbances from input–output data of multi-zone buildings. The unknown disturbance is a key challenge since it is not measurable but non-negligible. In this paper, we first present a principled method to aggregate a multi-zone building model into a single zone model, and show the aggregation is not as trivial as it has been assumed in the prior art. We then provide a method to identify the parameters of the model and the unknown disturbance for this aggregate (single-zone) model. Finally, we test our proposed identification algorithm to data collected from a multi-zone building testbed in Oak Ridge National Laboratory. A key insight provided by the aggregation method allows us to recognize under what conditions the estimation of the disturbance signal will be necessarily poor and uncertain, even in the case of a specially designed test in which the disturbances affecting each zone are known (as the case of our experimental testbed). This insight is used to provide a heuristic that can be used to assess when the identification results are likely to have high or low accuracy.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗