Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Geographic region”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A consistent dataset for the net income distribution for 190 countries and aggregated to 32 geographical regions from 1958 to 2015

Abstract. Data on income distributions within and across countries are becoming increasingly important for informing analysis of income inequality and understanding the distributional consequences of climate change. While datasets on income distribution collected from household surveys are available for multiple countries, these datasets often do not represent the same concept of inequality (or income concept) and therefore make comparisons across countries, over time and across datasets difficult. Here, we present a consistent dataset of income distributions across 190 countries from 1958 to 2015 measured in terms of net income. We complement the observed values in this dataset with values imputed from a summary measure of the income distribution, specifically the Gini coefficient. For the imputation, we use a recently developed nonparametric principal-component-based approach that shows an excellent fit to data on income distributions compared to other approaches. We also present another version of this dataset aggregated from the country level to 32 geographical regions. Our dataset is developed for the purpose of calibrating models such as integrated human–Earth system models with detailed data on income distributions. This dataset will enable more robust analysis of income distribution at multiple scales. The latest version of our data are available on Zenodo: https://doi.org/10.5281/zenodo.7093997 (Narayan et al., 2022b).

97 MATHEMATICS AND COMPUTING↗

Towards POI-based large-scale land use modeling: spatial scale, semantic granularity, and geographic context

The combination of spatial distribution, semantic characteristics, and sometimes temporal dynamics of POIs inside a geographic region can capture its unique land use characteristics. Most previous studies on POI-based land use modeling research focused on one geographic region and select one spatial scale and semantic granularity for land use characterization. There is a lack of understanding on the impact of spatial scale, semantic granularity, and geographic context on POI-based land use modeling, particularly large-scale land use modeling. In this study, we developed a scalable POI-based land use modeling framework and examined the impact of these three factors on POI-based land use characterization using data from three geographic regions. We developed a unified semantic representation framework for POI semantics that can help fuse heterogeneous POI data sources. Then, by combining POIs with a neural network language model, we developed a spatially explicit approach to learn the embedding representation of POIs and AOIs. We trained multiple supervised classifiers using AOI embeddings as input features to predict AOI land use at different semantic granularities. The classification performance of different land use classes was analyzed and compared across three geographic regions to identify the semantic representativeness of POI-based AOI embedding and the impact of geographic context.

58 GEOSCIENCES↗

Location Identifiers, Metadata, and Map for Field Measurements at the East-Taylor Watershed Community Observatory, Colorado, USA (Version 3.3)

This dataset contains identifiers, metadata, and a map of the locations where field measurements have been conducted at the East-Taylor Watershed Community Observatory located in the Upper Colorado River Basin, United States. This is version 3.3 of the dataset and replaces the prior version 3.2 (see below for details on changes between the versions). Dataset description: The East River-Taylor Watershed is the primary field site of the Watershed Function Scientific Focus Area (WFSFA) and the Rocky Mountain Biological Laboratory. Researchers from several institutions generate highly diverse hydrological, biogeochemical, climate, vegetation, geological, remote sensing, and model data at the East-Taylor Watershed in collaboration with the WFSFA. Thus, the purpose of this dataset is to maintain an inventory of the field locations and instrumentation to provide information on the field activities in the East-Taylor Watershed and coordinate data collected across different locations, researchers, and institutions. The dataset contains (1) a README file with information on the various files, (2) three csv files describing the metadata collected for each surface point location, plot and region registered with the WFSFA, (3) csv files with metadata and contact information for each surface point location registered with the WFSFA, (4) a csv file with with metadata and contact information for plots, (5) a csv file with metadata for geographic regions and sub-regions within the watershed, (6) a compiled xlsx file with all the data and metadata which can be opened in Microsoft Excel, (7) a kml map of the locations plotted in the watershed which can be opened in Google Earth, (8) a jpg image of the kml map which can be viewed in any photo viewer, and (9) a zipped file with the registration templates used by the SFA team to collect location metadata. The zipped template file contains two csv files with the blank templates (point and plot), two csv files with instructions for filling out the location templates, and one compiled xlsx file with the instructions and blank templates together. Additionally, the templates in the xlsx include drop down validation for any controlled metadata fields. Persistent location identifiers (Location_ID) are determined by the WFSFA data management team and are used to track data and samples across locations. Dataset uses: This location metadata is used to update the Watershed SFA’s publicly accessible Field Information Portal (an interactive field sampling metadata exploration tool; https://wfsfa-data.lbl.gov/watershed/), the kml map file included in this dataset, and other data management tools internal to the Watershed SFA team. Version Information: The latest version of this dataset publication is version 3.3. This version contains 167 new point locations, 1 new plot, and 2 new geographic regions. Overall, there are a total of 1439 point locations, 75 plots, and 54 geographic regions. Additionally, the kml map of locations and image now includes two boundaries (Upper Ohio Creek (UO) and Carbon Creek (CA)) outside of the East River watershed (USGS HUC-10) and accompanying stream network that represents areas of focus. Refer to methods for further details on the version history. This dataset will be updated on a periodic basis with new measurement location information. Researchers interested in having their East-Taylor Watershed measurement locations added to this list should reach out to the WFSFA data management team at wfsfa-data@googlegroups.com. Acknowledgments: Please cite this dataset if using any of the location metadata in other publications or derived products. If using the location metadata for the 2018 NEON hyperspectral campaign, additionally cite Chadwick et al. (2020). doi:10.15485/1618130. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

2018 NEON and 2025 CHESS Campaigns↗

Denoising Seismograms in the Time Domain Using a Deep Learning Model

Deep learning has emerged as a transformative tool for enhancing the extraction of reliable information from seismograms, addressing the increasing demand for precise and efficient seismic data analysis. We introduce an innovative encoder–decoder deep learning model, named WaveDenoiser, designed for noise reduction in the time domain, thereby eliminating the need for spectrogram computations that have been used for existing deep learning tools and significantly improving processing speed. Utilizing the benchmark dataset that is Stanford Earthquake Dataset, we developed three models of varying sizes: base, medium, and large. Notably, the large (referred to as WaveDenoiser) model demonstrated superior performance, achieving a median signal‐to‐noise ratio improvement of 8.8 dB on in‐distribution unseen data (in the same geographic region) and 7.7 dB on out‐distribution unseen data (in a new geographic region), outpacing both the base and medium models. Further evaluation of the WaveDenoiser model revealed a reduction in median arrival‐time errors by 0.02 s for P waves and 0.01 s for S waves when processing waveforms prior to phase picking using PhaseNet on in‐distribution unseen data. When tested on out‐distribution unseen data, the model also effectively reduced the P‐wave median arrival‐time error by 0.02 and 0.01 s in median arrival‐time error for S waves. Importantly, the application of WaveDenoiser resulted in a significant reduction of phase picking outliers by 1.1% to 3.6% for both P and S waves. In addition, we achieved over five times acceleration in processing speed compared with the seisBench implementation of DeepDenoiser. Our findings underscore the potential of WaveDenoiser as a powerful tool for improving seismic data analysis and processing efficiency.

P-waves↗

An OpenStreetMaps based tool to study the energy demand and emissions impact of electrification of medium and heavy-duty freight trucks

In this paper, we present the mathematical formulation of an OpenStreetMaps (OSM) based tool that compares the costs and emissions of long-haul medium and heavy-duty (M&HD) electric and diesel freight trucks, and determines the spatial distribution of added energy demand due to M&HD EVs. The optimization utilizes a combination of information on routes from OSM, utility rate design data across the United States, and freight volume data, to determine these values. In order to deal with the computational complexity of this problem, we formulate the problem as a convex optimization problem that is scalable to a large geographic area. In our analysis, we further evaluate various scenarios of utility rate design (energy charges) and EV penetration rate across different geographic regions and their impact on the operating cost and emissions of the freight trucks. Our approach determines the net emissions reduction benefits of freight electrification by considering the primary energy source in different regions. Such analysis will provide insights to policy makers in designing utility rates for electric vehicle supply equipment (EVSE) operators depending upon the specific geographic region and to electric utilities in deciding infrastructure upgrades based on the spatial distribution of the added energy demand of M&HD EVs. To showcase the results, a case study for the U.S. state of Texas is conducted.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fusing time-varying mosquito data and continuous mosquito population dynamics models

Climate change is arguably one of the most pressing issues affecting the world today and requires the fusion of disparate data streams to accurately model its impacts. Mosquito populations respond to temperature and precipitation in a nonlinear way, making predicting climate impacts on mosquito-borne diseases an ongoing challenge. Data-driven approaches for accurately modeling mosquito populations are needed for predicting mosquito-borne disease risk under climate change scenarios. Many current models for disease transmission are continuous and autonomous, while mosquito data is discrete and varies both within and between seasons. This study uses an optimization framework to fit a non-autonomous logistic model with periodic net growth rate and carrying capacity parameters for 15 years of daily mosquito time-series data from the Greater Toronto Area of Canada. The resulting parameters accurately capture the inter-annual and intra-seasonal variability of mosquito populations within a single geographic region, and a variance-based sensitivity analysis highlights the influence each parameter has on the peak magnitude and timing of the mosquito season. This method can easily extend to other geographic regions and be integrated into a larger disease transmission model. This method addresses the ongoing challenges of data and model fusion by serving as a link between discrete time-series data and continuous differential equations for mosquito-borne epidemiology models.

97 MATHEMATICS AND COMPUTING↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA's Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI's footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI's waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

54 ENVIRONMENTAL SCIENCES↗

Robustness of the Stochastic Parameterization of Subgrid-Scale Wind Variability in Sea Surface Fluxes

Abstract High-resolution numerical models have been used to develop statistical models of the enhancement of sea surface fluxes resulting from spatial variability of sea surface wind. In particular, studies have shown that flux enhancement is not a deterministic function of the resolved state. Previous studies focused on single geographical areas or used a single high-resolution numerical model. This study extends the development of such statistical models by considering six different high-resolution models, four different geographical regions, and three different 10-day periods, allowing for a systematic investigation of the robustness of both the deterministic and stochastic parts of the data-driven parameterization. Results indicate that the deterministic part, based on regressing the unresolved normalized flux onto resolved-scale normalized flux and precipitation, is broadly robust across different models, regions, and time periods. The statistical features of the stochastic part of the model (spatial and temporal autocorrelation and parameters of a Gaussian process fit to the regression residual) are also found to be robust and not strongly sensitive to the underlying model, modeled geographical region, or time period studied. Best-fit Gaussian process parameters display robust spatial heterogeneity across models, indicating potential for improvements to the statistical model. These results illustrate the potential for the development of a generic, explicitly stochastic parameterization of sea surface flux enhancements dependent on wind variability.

Endo, Kota↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

The impact of adjusting for hysterectomy prevalence on cervical cancer incidence rates and trends among women aged 30 years or older—United States, 2001-2019

Abstract Hysterectomy protects against cervical cancer when the cervix is removed. However, measures of cervical cancer incidence often fail to exclude women with a hysterectomy from the population-at-risk denominator, underestimating and distorting disease burden. In this study, we estimated hysterectomy prevalence from the Behavioral Risk Factor Surveillance System surveys to remove the women who were not at risk of cervical cancer from the denominator and combined these estimates with the US Cancer Statistics data. From these data, we calculated age-specific and age-standardized incidence rates for women aged >30 years from 2001-2019, adjusted for hysterectomy prevalence. We calculated the difference between unadjusted and adjusted incidence rates and examined trends by histology, age, race and ethnicity, and geographic region using joinpoint regression. The hysterectomy-adjusted cervical cancer incidence rate from 2001-2019 was 16.7 per 100 000 women—34.6% higher than the unadjusted rate. After adjustment, incidence rates were higher by approximately 55% among Black women, 56% among those living in the East South Central division, and 90% among women aged 70-79 and ≥80 years. These findings underscore the importance of adjusting for hysterectomy prevalence to avoid underestimating cervical cancer incidence rates and masking disparities by age, race, and geographic region. This article is part of a Special Collection on Gynecological Cancers.

Public, Environmental & Occupational Health↗

Simulating the Impact of Dynamic Rerouting on Metropolitan-scale Traffic Systems

The rapid introduction of mobile navigation aides that use real-time road network information to suggest alternate routes to drivers is making it more difficult for researchers and government transportation agencies to understand and predict the dynamics of congested transportation systems. Computer simulation is a key capability for these organizations to analyze hypothetical scenarios; however, the complexity of transportation systems makes it challenging for them to simulate very large geographical regions, such as multi-city metropolitan areas. In this article, we describe enhancements to the Mobiliti parallel traffic simulator to model dynamic rerouting behavior with the addition of vehicle controller actors and vehicle-to-controller reroute requests. The simulator is designed to support distributed-memory parallel execution using discrete event simulation and be scalable on high-performance computing platforms. We demonstrate the potential of the simulator by analyzing the impact of varying the population penetration rate of dynamic rerouting on the San Francisco Bay Area road network. Using high-performance parallel computing, we can simulate a day in the San Francisco Bay Area with 19 million vehicle trips with 50 percent dynamic rerouting penetration over a road network with 0.5 million nodes and 1 million links in less than three minutes. We present a sensitivity study on the dynamic rerouting parameters, discuss the simulator’s parallel scalability, and analyze system-level impacts of changing the dynamic rerouting penetration. Furthermore, we examine the varying effects on different functional classes and geographical regions and present a validation of the simulation results compared to real-world data.

97 MATHEMATICS AND COMPUTING↗

A Data Processing Pipeline To Extract A Knowledge Graph From Sec Documents For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.

Weaver, GabrielA.↗

Evaluation of Extreme Weather Impacts on Utility-scale Photovoltaic Plant Performance in the United States

The global energy system is undergoing significant changes, including a shift in energy generating technologies to more renewable energy sources. However, the dependence of renewable energy sources on local environmental conditions could also increase disruptions in service through exposures to compound, extreme weather events. By fusing three diverse datasets (operations and maintenance tickets, weather data, and production data), this analysis presents a novel methodology to identify and evaluate performance impacts arising from extreme weather events across diverse geographical regions. Text analysis of maintenance tickets identified snow, hurricanes, and storms as the leading extreme weather events affecting photovoltaic plants in the United States. Statistical techniques and machine learning were then implemented to identify the magnitude and variability of these extreme weather impacts on site performance. Impacts varied between event and non-event days, with snow events causing the greatest reductions in performance (54.5%), followed by hurricanes (12.6%) and storms (1.1%). Machine learning analysis identified key features in determining if a day is categorized as low performing, such as low irradiance, geographic location, weather features, and site size. The analysis improves our understanding of compound, extreme weather event impacts on photovoltaic systems, which can inform planning activities, especially as the industry continues to expand into new geographic and climatic regions around the world.

14 SOLAR ENERGY↗

Rapid increase in tropospheric ozone over Southeast Asia attributed to changes in precursor emission source regions and sectors

Observations indicate that tropospheric ozone (O 3 ) concentrations over Southeast Asia have been increasing rapidly since the 1990s. Here, we quantify source contributions from geographical regions and emission sectors of the two distinct types of O 3 precursors, i.e., nitrogen oxides (NO x ) and volatile organic compounds (VOCs), to the increase in tropospheric O 3 in Southeast Asia during 1990–2019 using an O 3 source tagging technique implemented in a global chemistry-climate model. In this work, the results show that although local anthropogenic emission of NO x in Southeast Asia only contributes 18% of the annual averaged near-surface O 3 concentration, the increase in local NO x emission dominates the increasing trend of O 3 concentration in Southeast Asia, accounting for 107% of the regional averaged trend of 1.07 ppb decade -1 . Increases in NO x emissions from East Asia and South Asia explain 29% of the increasing trend, but 9% is offset by the emission reduction in North America. Ground transportation is responsible for 79% of the rapid O 3 increase, followed by 39% contribution from international shipping. Because an increase in anthropogenic NO x emissions enhances the O 3 production efficiency by VOCs, the increase in near-surface O 3 concentrations in Southeast Asia is thereby largely contributed by methane and biogenic VOCs.

54 ENVIRONMENTAL SCIENCES↗

Evaluating Direct and Indirect Influence on EV Charging Stations Across the US

The adoption of new technology for electric vehicles (EV) and mobility applications can bring underappreciated vulnerabilities to the power grid. One area of potential fraud and adversarial influence is through the business ecosystem of startups that own and deploy EV technology. Yet, there are no models or analyses that map the network of organizations and people that have direct and indirect influence over technologies currently deployed in the grid. To fill this gap, we develop a multilayer network model to measure direct and indirect influence on EV charging stations. First, we create and adversarial socio-technical network (ASTN) model via a data fusion pipeline for different US regions of interest (ROI). Then, we develop an integrated ASTN for Chicago, Los Angeles, New York, and Philadelphia. We rank EV charging companies direct influence within each geographic region as well as indirect influence via social network analysis. While some companies have strong direct and indirect influence (i.e., ChargePoint) others show a mismatch between their influence over charging stations and their position within the social network. For example, Tesla has strong direct influence on stations and weak indirect influence over competitors. In contrast, 7Charge has weak direct influence over stations, but strong indirect influence over competitors.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

Empirical Validation of UBEM: An Assessment of Bias in Urban Building Energy Modeling for Chicago

Residential and commercial buildings currently account for 30% of total global final energy consumption. Urban-scale building energy modeling (UBEM) can enable scalable investments and unlock building improvements by quantifying energy, demand, emissions, and cost reductions of specific measures or packages for building-specific technologies in large geographic regions. While the sophistication of UBEM data sources and technologies have increased dramatically in the past decade, there remains a knowledge gap for empirical validation and sources of bias between building-specific energy models and measured data at varying geographic scales.As UBEM continues to develop, systemic analysis of accuracy, bias, and limitations of the resulting models is necessary to inform best practices and move toward standardization. These are characterized for the Automatic Building Energy Modeling (AutoBEM) software suite with an initial case study involving metered electricity consumption data from 247,188 buildings in Chicago, Illinois, USA - averaged across years 2019-2021 - compared to the following datasets: (1) the AutoBEM-generated nation-scale Model America version 2 (MAv2) data for 596,064 buildings, (2) tax assessor data for 579,829 buildings, (3) tax assessor data filled with MAv2, and (4) 102 representative dynamic archetypes. The accuracy is reported for every building type and vintage combination, along with multiple sources of bias for unique building descriptors. The AutoBEM simulation workflow produced energy consumption estimates that closely match aggregated metered electricity consumption data for different types of buildings constructed during various time periods at the city scale - with initial normalized mean bias error of 10.9%, and 1.1% after removing outliers. Contribution of statistically significant factors including building type, land use, age, and size to variance in UBEM bias is quantified.

Garg, Ankur↗

CareWELL: Multimodal Region Representation Learning with Spatial Contexts for Urban Health

Rapid urbanization affects living environments by intensifying exposure to air pollution, heat, noise, and urban dynamics, which together contribute to uneven health outcomes across neighborhoods. For instance, cardiovascular, respiratory, and mental health conditions are each influenced by distinct exposures such as air pollution, extreme temperatures, or limited access to green space. These heterogeneous patterns require understanding the characteristics of geographic regions in order to explain why urban health risks vary across urban areas. Recent work in self-supervised region representation learning provides a promising way to model such characteristics from multimodal geospatial data. However, existing methods face two major limitations: (i) they often depend on non-public datasets, limiting reproducibility and applicability, and (ii) their generic pretraining objectives overlook health-relevant determinants, including temporal variability in environmental exposures and inequalities in social conditions. To address these gaps, we propose Context-Aware Region rEpresentation with Weather, Environment, and Location Learning (CareWELL). CareWELL leverages large language models to encode seasonal variability in weather, employs contrastive learning to align geo-coordinate and weather representations, and introduces a context-aware objective that integrates socio-demographic factors while preserving spatial correlations. We evaluate CareWELL by predicting six urban health outcomes in Manhattan, New York City, and demonstrate that CareWELL consistently outperforms state-of-the-art baselines as well as a traditional spatial computing method. These results suggest the importance of context-aware pretraining objectives for learning health-relevant region representations.

Namgung, Min [ORNL]↗