Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “geospatial database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Machine-Learning-Based Mapping and Modeling of Solar Energy with Ultra-High Spatiotemporal Granularity

Despite the rapid growth of solar energy, we still lack a dynamic, high-fidelity database that tracks the spatiotemporal variations of solar PVs and their associated infrastructures across different places at a spatially resolved scale. The absence of such data presents a barrier to various applications such as solar PV growth projection, solar energy integration, solar incentive design, and climate risk assessment. In this project, we aim to bridge this gap by developing AI-based algorithms to extract granular information about solar PV installations and their associated infrastructures (i.e., distribution grids) from widely available unstructured data like remote sensing images and street views. As a result, we have built the Solar Energy Atlas, a fine-grained, large-scale geospatial overlay of distributed solar PVs and distribution grids. On top of it, we have advanced the understanding of solar adoption and distribution grid vulnerability to climate-induced extremes. Our major contributions can be summarized as follow: (1) By developing new AI algorithms, we have built the most comprehensive solar PV spatiotemporal database covering the entire US. This is the first time we obtained the exact GPS locations, size, subtype, and installation year information for rooftop solar PVs across the US. This database can be used for solar PV growth projection, solar energy integration, solar energy policy analysis and design, and spatially-resolved climate risk assessment. (2) Leveraging this database, we have uncovered the socioeconomic driving factors that are correlated with earlier onset of solar adoption and higher saturated adoption levels. We have identified the heterogeneity in the effects of different types of financial incentives on solar adoption and provided implications for tailoring incentive design based on local income levels to promote equitable solar adoption. (3) We have developed a distribution grid GIS mapping algorithm which can obtain granular geospatial and topology information about distribution grids using multi-modal open data, reducing the dependency on hard-to-obtain smart meter data of conventional approaches. It shows effectiveness in both the U.S. and Sub-Saharan Africa. Using this algorithm, we have uncovered the non-uniform vulnerability of distribution grids to wildfires in California in the aspects of undergrounding protection and Distributed Energy Resources (DER) preparedness. This has provided important implications for improving the affordability and equity of grid adaptation approaches. (3) We have made our produced database publicly available and provided user-friendly interface to enable various stakeholders and the general public to interact with the data. We have also integrated the produced data into the Data Commons platform to enable the public to access the data and correlate it with other location-specific characteristics simply using natural language as queries. The impact of our project is three-fold: (1) New algorithms for mapping solar PVs and distribution grids across space and time, which are open source to facilitate researchers and industry; (2) New databases of solar PVs and distribution grids that have been made publicly available for engineering, social, and policy applications; (3) New understandings and actionable insights on the potential approaches to promoting solar adoption and reducing energy infrastructure vulnerabilities. In this report, we start by discussing the project background and motivation (section 5), followed by the overview of project objectives (section 6). Results and discussion for each task are presented in section 7. Significant accomplishments are summarized in section 8. This report will be concluded by discussing the paths forwards (section 9), products (section 10), and team roles (section 11).

14 SOLAR ENERGY↗

A global metagenomic map of urban microbiomes and antimicrobial resistance

We present a global atlas of 4,728 metagenomic samples from mass-transit systems in 60 cities over 3 years, representing the first systematic, worldwide catalog of the urban microbial ecosystem. This atlas provides an annotated, geospatial profile of microbial strains, functional characteristics, antimicrobial resistance (AMR) markers, and genetic elements, including 10,928 viruses, 1,302 bacteria, 2 archaea, and 838,532 CRISPR arrays not found in reference databases. We identified 4,246 known species of urban microorganisms and a consistent set of 31 species found in 97% of samples that were distinct from human commensal organisms. Profiles of AMR genes varied widely in type and density across cities. Cities showed distinct microbial taxonomic signatures that were driven by climate and geographic differences. These results constitute a high-resolution global metagenomic atlas that enables discovery of organisms and genes, highlights potential public health and forensic applications, and provides a culture-independent view of AMR burden in cities.

59 BASIC BIOLOGICAL SCIENCES↗

Appendices for Geothermal Exploration Artificial Intelligence Report

The Geothermal Exploration Artificial Intelligence looks to use machine learning to spot geothermal identifiers from land maps. This is done to remotely detect geothermal sites for the purpose of energy uses. Such uses include enhanced geothermal system (EGS) applications, especially regarding finding locations for viable EGS sites. This submission includes the appendices and reports formerly attached to the Geothermal Exploration Artificial Intelligence Quarterly and Final Reports. The appendices below include methodologies, results, and some data regarding what was used to train the Geothermal Exploration AI. The methodology reports explain how specific anomaly detection modes were selected for use with the Geo Exploration AI. This also includes how the detection mode is useful for finding geothermal sites. Some methodology reports also include small amounts of code. Results from these reports explain the accuracy of methods used for the selected sites (Brady Desert Peak and Salton Sea). Data from these detection modes can be found in some of the reports, such as the Mineral Markers Maps, but most of the raw data is included the DOE Database which includes Brady, Desert Peak, and Salton Sea Geothermal Sites.

15 GEOTHERMAL ENERGY↗

LiAISON (Life-cycle Assessment Integration into Scalable Open-source Numerical models) [SWR-24-01]

We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE)7. We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Under baseline projections (i.e., no decarbonization goals), neither process reaches parity with the incumbent technology across several environmental metrics. Under the decarbonization scenarios, the underlying sectoral shifts result in declining impacts over time, compared to 2020 levels, except for metal depletion levels, which increase. The background shifts postulate a heavily decarbonized economy and energy system, which help technologies reach parity with SMR between 2040-2050 (RCP2.6) and 2030-2040 (RCP1.9) for global warming. Despite declines across several other metrics over time, neither PtH2 technology break even with SMR by 2100 besides for global warming. Scientific publication available here: https://pubs.acs.org/doi/full/10.1021/acs.est.2c04246

Ghosh, Tapajyoti↗

Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) for Prospective Impact Analysis of Novel Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 Degrees Celsius or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics. The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Additionally we compare our results by linking two other prospective models with LiAISON - GCAM (Global Change Assessment Model) and ReEDS (Regional Energy Deployment System) to analyze the effect of changing background scenarios using varying predictions in life cycle analysis.

decarbonizing↗

Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) for Analyzing Emerging Low-Carbon Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 degrees Celsius or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics. The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Under baseline projections (i.e., no decarbonization goals), neither process reaches parity with the incumbent technology across several environmental metrics. Under the decarbonization scenarios, the underlying sectoral shifts result in declining impacts over time, compared to 2020 levels, except for metal depletion levels, which increase. The background shifts postulate a heavily decarbonized economy and energy system, which help technologies reach parity with SMR between 2040-2050 (RCP2.6) and 2030-2040 (RCP1.9) for global warming. Despite declines across several other metrics over time, neither PtH2 technology break even with SMR by 2100 besides for global warming.

decarbonizing↗

Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) for Analyzing Emerging Low-Carbon Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 degrees C or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Under baseline projections (i.e., no decarbonization goals), neither process reaches parity with the incumbent technology across several environmental metrics. Under the decarbonization scenarios, the underlying sectoral shifts result in declining impacts over time, compared to 2020 levels, except for metal depletion levels, which increase. The background shifts postulate a heavily decarbonized economy and energy system, which help technologies reach parity with SMR between 2040-2050 (RCP2.6) and 2030-2040 (RCP1.9) for global warming. Despite declines across several other metrics over time, neither PtH2 technology break even with SMR by 2100 besides for global warming.

decarbonizing↗

Towards Prospective LCA Using Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) Framework for Analyzing Emerging Low-Carbon Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 Degrees Celsius or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Additionally we compare our results by linking two other prospective models with LiAISON - GCAM(Global Change Assessment Model) and ReEDS (Regional Energy Deployment System) to analyze the effect of changing background scenarios using varying predictions in life cycle analysis.

emissions↗

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds Developed lands Areas >0.8 km (0.5 miles) from developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall↗

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall↗

IM3 Projected US Data Center Locations

IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall (ORCID:0000000328077088)↗

Carbon Transport and Storage Planning and Viability Support Tools

The EDX disCO2ver Carbon Transport and Storage Planning and Viability Support Tools are made up of the Carbon Storage Planning Inquiry Tool (CS PlanIT, Justman et al. 2024) and the Carbon Storage Technical Viability Approach Support Tool (CS TVA). Together, these tools support data access to support understanding data availability to support planning efforts for carbon transport and storage. The Carbon Storage Planning Inquiry Tool (CS PlanIT) is an online web mapping application designed to help users explore, query, and evaluate multiple data layers to support and accelerate carbon storage resource and feasibility assessments and planning efforts. CS PlanIT currently contains a range of datasets associated with geologic, technical, and infrastructure factors. The data sets can be filtered geographically for an area of interest to update statistics and charts within the dashboard. The dashboard is divided into different sections called widgets, relating to different steps in the carbon storage planning process. The resources in this submission include a link to PlanIT, as well as a data catalog and link to user documentation. The original citation for the CS PlanIT tool, which has now been integrated into the toolset here, was: - Devin Justman, Scott Pantaleone, Maneesh Sharma, Lucy Romeo, Paige Morkner, CS PlanIT (Carbon Storage Planning Inquiry Tool) , 6/28/2024, https://edx.netl.doe.gov/dataset/cs-planit-carbon-storage-planning-inquiry-tool, DOI: 10.18141/2377953 The Carbon Storage Technical Viability Approach Support (CS TVA) Tool displays spatial data availability for the many components of Geologic Carbon Storage (GCS) technical viability assessment (Creason al 2025). Identifying sites suitable for GCS requires evaluating the intersection of myriad factors, including reservoir conditions, subsurface and surface hazards, infrastructure requirements, and energy community metrics. The technical viability of a site can only be confirmed for instances where all these factors have data available, and where those data support viability. Additional Resources related to the Technical Viability Assessment Tool: - Julia Mulhern, Casey White, Araceli Lara, Neyda Cordero Rodriguez, Zachary Jackson, Jacob Shay, Gabriel Creason, MacKenzie Mark-Moser, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Database, 3/26/2025, https://edx.netl.doe.gov/dataset/edx4ccs-carbon-storage-technical-viability-approach-database , DOI:10.18141/1984655 - Julia Mulhern, MacKenzie Mark-Moser, Gabriel Creason, Casey White, Araceli Lara, Neyda Cordero Rodriguez, Zach Jackson, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Matrix, 3/27/2025, https://edx.netl.doe.gov/dataset/carbon-storage-technical-viability-approach-cs-tva-matrix , DOI: 10.18141/2539979 - Gabriel Creason, Zach Jackson, Neyda Cordero Rodriguez, Julia Mulhern, Casey White, Araceli Lara, MacKenzie Mark-Moser, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Data Availability Results Database, 3/27/2025, https://edx.netl.doe.gov/dataset/carbon-storage-technical-viability-approach-cs-tva-data-availability-results-database, DOI:10.18141/2538557

Carbon storage↗

IM3 Projected US Data Center Locations

IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall (ORCID:0000000328077088)↗

Development and Preliminary Analysis of a U.S. Geothermal Heat Pump Installation Database

This paper seeks to addresses the significant gap in the literature regarding the installation and adoption of geothermal heat pump (GHP) systems in the United States. While the "2021 U.S. Geothermal Power Production and District Heating Market Report" published by the National Renewable Energy Laboratory (NREL) focused on direct-use geothermal district heating systems, it did not include an analysis of GHP installations (Robins et al. 2021). To bridge this gap, NREL has compiled a novel database currently containing 70,470 records of GHP installations, primarily sourced from state well permits and small-scale studies. Our methodology emphasizes the collection, cleaning, and standardization of data, addressing challenges such as inconsistent reporting formats and privacy concerns. Despite limitations in data on capacity, costs, and performance, our preliminary geospatial analysis reveals insights into the distribution of GHP systems across urban and rural areas and climate zones. The paper highlights the importance of publicly accessible data for advancing GHP technology adoption with a discussion of existing data sources and their limitations, advocating for improved collaboration between NREL and industry stakeholders.

data collection↗

Locating Undocumented Wells Using Historical Oil and Gas Exploration Maps: A Case Study in Osage County, Oklahoma

Undocumented oil and gas wells lack reliable information about their locations and characteristics, making them difficult to identify. These wells can result in unanticipated delays and costs in the development of nearby surface and subsurface resources, and, if improperly plugged, can cause contamination. This study leverages historical petroleum exploration maps to locate such wells, focusing on Osage County, Oklahoma. Two sets of early 20th century oil and gas exploration maps by the United States Geological Survey were georeferenced and analyzed using a computer vision model to detect well symbols. The locations of detected wells were compared to the location of known wells in the database from the Bureau of Indian Affairs Osage Agency to identify potential undocumented wells. The analysis yielded over 500 potential undocumented wells, with dry holes constituting the largest fraction. Field verification confirmed the presence of some undocumented wells. Comparison with prior work revealed limited overlap, underscoring the complementary value of historical oil and gas maps for locating undocumented wells. This approach demonstrates the utility of integrating historical cartographic resources with modern geospatial and machine learning techniques to improve the identification and management of undocumented wells.

Energy - Petroleum↗

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

A Data Processing Pipeline for Adversarial Socio-Technical Network Analysis

With the rapid adoption of emerging technologies, there is a need to catalog and model sociotechnical interdependencies that have been historically used to influence the operation of Critical Infrastructure networks including the impacts of mergers and acquisitions, hostile takeovers, and foreign investment. Our research intends to address this need with two primary contributions. First, we have developed a data curation and processing pipeline to generate sociotechnical networks extracted from a variety of data sources including SEC filings and infrastructure asset databases. The pipeline, implemented in Apache Airflow, extracts and normalizes the representation of entities and relations, specified within ontologies. Our intent is to provide an extensible, machine-actionable approach to quickly communicate such models, reproduce previous results, and adapt them to new, unanticipated situations. Second, networks produced by our pipeline enable the development of graph-theoretic metrics that consider the properties of network components in addition to its topology. Metadata associated with network components---whether semantic, temporal, or geospatial---affects the alignment of generated networks with assumptions underlying complexity metrics. Validation of generated networks relative to component types defined by an ontology, may allow the research community to adapt metrics to the semantics of the domains being studied. Generated networks may be processed as knowledge, dynamic, or spatial graphs and enables a variety of analyses including automated reasoning and measures of network complexity. Automated reasoning views extracted entities and relations as a knowledge graph; this enables application of inference rules that represent historically-attested adversarial business methods and applies that behavior to a specific geographic context. Measures of network complexity, including degree distribution, reachability analyses, temporal analysis, and community detection can be adapted to indicate adversarial organizational influence.

97 MATHEMATICS AND COMPUTING↗