Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “U.S. Census”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Data items will occasionally be removed from OSM if they are misidentified, if they no longer exist, if they are duplicates of another item, or similar. For that reason, updated versions of this database may not contain all data center locations included in previous versions. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Revised monthly energy generation estimates for 1,500 hydroelectric power plants in the United States

Abstract The U.S. Energy Information Administration (EIA) conducts a regular survey (form EIA-923) to collect annual and monthly net generation for more than ten thousand U.S. power plants. Approximately 90% of the ~1,500 hydroelectric plants included in this data release are surveyed at annual resolution only and thus lack actual observations of monthly generation. For each of these plants, EIA imputes monthly generation values using the combined monthly generating pattern of other hydropower plants within the corresponding census division. The imputation method neglects local hydrology and reservoir operations, rendering the monthly data unsuitable for various research applications. Here we present an alternative approach to disaggregate each unobserved plant’s reported annual generation using proxies of monthly generation—namely historical monthly reservoir releases and average river discharge rates recorded downstream of each dam. Evaluation of the new dataset demonstrates substantial and robust improvement over the current imputation method, particularly if reservoir release data are available. The new dataset—named RectifHyd—provides an alternative to EIA-923 for U.S. scale, plant-level, monthly hydropower net generation (2001–2020). RectifHyd may be used to support power system studies or analyze within-year hydropower generation behavior at various spatial scales.

13 HYDRO ENERGY↗

Household Energy Efficiency Analysis in San Jose, California

Many households in the City of San Jose, California, could save hundreds of dollars annually on their energy bills and reduce carbon emissions with energy efficiency retrofits and upgrades in their homes and apartments. As part of the U.S. Department of Energy's (DOE) Communities LEAP (Local Energy Action Program) pilot, the National Renewable Energy Laboratory (NREL) analyzed energy efficiency and electrification upgrades for 83,000 housing units in low-to-moderate-income census tracts experiencing disproportionately high energy burden in San Jose. The results in this fact sheet only pertain to the housing units included in the analysis.

Communities LEAP↗

Alternative communication network designs for an operational Plato 4 CAI system

The cost of alternative communications networks for the dissemination of PLATO IV computer-aided instruction (CAI) was studied. Four communication techniques are compared: leased telephone lines, satellite communication, UHF TV, and low-power microwave radio. For each network design, costs per student contact hour are computed. These costs are derived as functions of student population density, a parameter which can be calculated from census data for one potential market for CAI, the public primary and secondary schools. Calculating costs in this way allows one to determine which of the four communications alternatives can serve this market least expensively for any given area in the U.S. The analysis indicates that radio distribution techniques are cost optimum over a wide range of conditions.

Mobley, R. E., Jr.↗

Adapting Agrivoltaics for Solar Mini-Grids in Haiti

With less than 2% of the rural population with access to electricity and almost half the population facing acute hunger, Haiti faces interconnected challenges of energy poverty and food insecurity. One solution to help address energy poverty in Haiti has been the development of distributed solar, particularly solar mini-grids. However, often the land best suited for deploying solar generation is also best suited for agriculture by smallholder farmers, thereby creating a potentially complicated tension between energy access and food security. To address this tension, solar developers, agricultural specialists, and researchers are jointly examining a novel solution called "agrivoltaics." Agrivoltaics is a shared land-use solution that is rapidly expanding in established solar markets like the United States, Europe, and Asia that pairs solar with agriculture, producing electricity and providing space for crops and animal grazing under and between panels. As part of the Energy Access Partnership for Haiti with the U.S. Agency for International Development (USAID), the National Renewable Energy Laboratory (NREL) performed an initial feasibility analysis and stakeholder engagement project to evaluate the potential for agrivoltaics in mini-grid contexts in Haiti. The analysis considered typical 100-kW and larger 1-MW mini-grids in towns across Haiti and developed two example agrivoltaic archetypes based on key local inputs, including solar irradiance, production data from the agricultural census, market prices, stakeholder interviews, and existing agrivoltaic research. See NREL/TP-7A40-89399 for the Haitian Creole translation of this document.

14 SOLAR ENERGY↗

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics↗

Prevalence of Americans reporting a family history of cancer indicative of increased cancer risk: Estimates from the 2015 National Health Interview Survey

The collection and evaluation of family health history in a clinical setting presents an opportunity to discuss cancer risk, tailor cancer screening recommendations, and identify people with an increased risk of carrying a pathogenic variant who may benefit from referral to genetic counseling and testing. National recommendations for breast and colorectal cancer screening indicate that men and women who have a first-degree relative affected with these types of cancers may benefit from talking to a healthcare provider about starting screening at an earlier age and other options for cancer prevention. The prevalence of reporting a first-degree relative who had cancer was assessed among adult respondents of the 2015 National Health Interview Survey who had never had cancer themselves (n = 27,999). We found 35.6% of adults reported having at least one first-degree relative with cancer at any site. Significant differences in reporting a family history of cancer were observed by sex, age, race/ethnicity, educational attainment, and census region. Nearly 5% of women under age 50 and 2.5% of adults under age 50 had at least one first-degree relative with breast cancer or colorectal cancer, respectively. We estimated that 5.8% of women had a family history of breast or ovarian cancer that may indicate increased genetic risk. A third of U.S. adults who have never had cancer report a family history of cancer in a first-degree relative. Furthermore, this finding underscores the importance of using family history to inform discussions about cancer risk and screening options between healthcare providers and their patients.

60 APPLIED LIFE SCIENCES↗

Activities of the US Geological Survey in Applications of Remote Sensing in the Chesapeake Bay Region

The application of remote sensing in the Chesapeake Bay region has been a central concern of three project activities of the U.S. Geological Survey: two are developmental, and one is operational. The two developmental activities were experiments in land-use and land-cover inventory and change detection using remotely sensed data from aircraft and from the LANDSAT and Skylab satellites. One of these is CARETS (Central Atlantic Regional Ecological Test Site). The other developmental task is the Census Cities Experiment in Urban Change Detection. The present major concern is an operational land-use and land-cover data-analysis program, including a supporting geographical information system.

Wray, J. R.↗

National CMM Insights Platform

The National CMM Insights Platform is an interactive application. The tool showcases the U.S. National CMM Insights Dataset. It is intended to be used for research and comparison purposes, see full disclaimer and credits. Description: The National CMM Insights Platform is an interactive tool. The tool showcases the National CMM Insight Dataset. Application link: https://arcgis.netl.doe.gov/portal/apps/experiencebuilder/experience/?id=745ba8726d894a7f86c5d1859116c009 For additional information, please check out the StoryMap documentation: National CMM Insights Platform Guide The National CMM Insights Platform and Dataset provides access to aggregated publicly available information pertaining to census tracts, counties and states.

AS↗

BioSiting Tool (BioSiting) v2

The BioSiting Tool provides a geospatial interface for analyzing bioeconomy resources and infrastructure across the continental U.S. The tool integrates empirical and modeled data from a broad range of sources. Bioeconomy resources mapped in the tool include agricultural residues, forest residues, municipal solid waste streams, food waste, manure, fats, oils and greases and potential yields of energy crops. Infrastructure mapped in the tool includes biorefineries, material recovery facilities, anaerobic digesters, wastewater treatment plants, combustion plants, district energy systems, crude oil pipelines, petroleum pipelines, natural gas pipelines, railways and freight terminals. Additional data layers include environmental justice indicators at the census tract level and carbon dioxide geologic storage potential. Users can select a location on the map, define a buffer radius in kilometers and generate an inventory of all bioecomony resources within the buffer zone. Data from the tool can be downloaded from individual buffer zones, or at the state or national level.

Huntington, Tyler↗

Heat Vulnerability Index Development and Mapping

Extreme heat is one of the leading causes of weather-related deaths in the U.S. Exposure to extreme heat will be exacerbated due to global climate change. It is thus crucial to design a key performance indicator, heat vulnerability index (HVI), to represent overall heat risk which can help identify susceptible regions and sub-populations in cities in the face of heatwaves. Most existing HVI tools only consider outdoor heat exposure. This paper developed an HVI web map tool incorporating both the outdoor and the indoor heat exposure, as well as population sensitivity and adaptation capability across census tracts in the city of Fresno, California. The tool can assist the planning of infrastructure and resources to reduce residents’ vulnerability to extreme heat events.

Xu, Yujie↗

Strategic planning for aircraft noise route impact analysis: A three dimensional approach

The strategic routing of aircraft through navigable and controlled airspace to minimize adverse noise impact over sensitive areas is critical in the proper management and planning of the U.S. based airport system. A major objective of this phase of research is to identify, inventory, characterize, and analyze the various environmental, land planning, and regulatory data bases, along with potential three dimensional software and hardware systems that can be potentially applied for an impact assessment of any existing or planned air route. There are eight data bases that have to be assembled and developed in order to develop three dimensional aircraft route impact methodology. These data bases which cover geographical information systems, sound metrics, land use, airspace operational control measures, federal regulations and advisories, census data, and environmental attributes have been examined and aggregated. A three dimensional format is necessary for planning, analyzing space and possible noise impact, and formulating potential resolutions. The need to develop this three dimensional approach is essential due to the finite capacity of airspace for managing and planning a route system, including airport facilities. It appears that these data bases can be integrated effectively into a strategic aircraft noise routing system which should be developed as soon as possible, as part of a proactive plan applied to our FAA controlled navigable airspace for the United States.

Bragdon, C. R.↗

GeoCricket

SAND2025-12229O Geospatial Critical Infrastructure and Census Data Stockpile Tool (GeoCricket) is a set of functions that collect critical infrastructure and census data for use in the Resilient Node Cluster Analysis Tool (ReNCAT) and Quantum Geographic Information System Social Burden Calculator. It can also act to inform other place-based work. The code queries public-facing Representational State Transfer (REST) servers to collect geospatial data related to a specific area. It then exports that data as standard geographic information system file types or as a .csv file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Haines, John↗

IRA Energy Community Data Layers

Data, geospatial data resources, and the linked mapping tool and web services reflect data for two types of potentially qualifying energy communities: 1) Census tracts and directly adjoining tracts that have had coal mine closures since 1999 or coal-fired electric generating unit retirements since 2009. These census tracts qualify as energy communities. 2) Metropolitan statistical areas (MSAs) and non-metropolitan statistical areas (non-MSAs) that are energy communities for 2023 and 2024, along with their fossil fuel employment (FFE) status. Additional information on energy communities and related tax credits can be accessed on the Interagency Working Group on Coal & Power Plant Communities & Economic Revitalization Energy Communities website (https://energycommunities.gov/energy-community-tax-credit-bonus/). Use limitations: these spatial data and mapping tool may not be relied upon by taxpayers to substantiate a tax return position or for determining whether certain penalties apply and will not be used by the IRS for examination purposes. The mapping tool does not reflect the application of the law to a specific taxpayer’s situation, and the applicable Internal Revenue Code provisions ultimately control.

Census Tract↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗