Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Population estimates”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Worldwide Population Estimates for Small Geographic Areas

This chapter discusses the basis for these population estimates, scope, and limitations based on experiences in development of massive global population datasets and usage of these datasets as a basis for sampling design. It presents tools and approaches for using these georeferenced population estimates for complex household survey sampling. The chapter provides a unified resource for understanding current gridded population datasets, their use in survey research, and promising areas of future work toward improved population estimates. It includes an overview of popular gridded population datasets, common methodological frameworks for developing gridded population estimates, and discusses pros and cons for selecting a gridded population dataset from a survey perspective. The chapter focuses on survey methods that use gridded population estimates specifically. It highlights differences between census population estimates and gridded population estimates that may be relevant when planning data collection. The chapter presents a case study of gridded population data and sampling methods in Nigeria.

Amer, Safaa↗

LandScan mosaic enables high-resolution gridded population estimates with explicit uncertainty

Gridded population datasets represent high-resolution distributions of human occupancy, enabling informed decision-making across a broad range of fields. These data products are valuable for assessing environmental risk, urban development, disaster preparedness and resource allocation—areas where accurate population estimates directly enhance policy effectiveness and optimize resource distribution. Despite the importance of gridded population datasets, traditional population modeling approaches often overlook inherent uncertainties in the estimation process. This limitation can create a false sense of certainty in population estimates, potentially leading to flawed decisions by those who rely on the data. To address this methodological gap, we introduce a probabilistic machine learning modeling framework, LandScan Mosaic, that explicitly incorporates uncertainty into the population modeling process. Our approach systematically quantifies uncertainty in three key modeling parameters of the LandScan HD gridded population dataset: building use types, floor counts, and occupancy rates. By employing Monte Carlo simulations, we propagate these uncertainties through the modeling process, yielding probability distributions of population counts in place of deterministic point estimates. We demonstrate the practical application of this framework in Iloilo City, Philippines, using structured decision-making techniques and our probabilistic estimates to identify and prioritize areas most affected by projected flooding, supporting targeted interventions that address both economic and social risks. In doing so, we propose a population-specific approach for incorporating confidence into structured decision making processes. Through a comparative analysis with conventional deterministic approaches and point estimate approaches, including LandScan HD and WorldPop, we evaluate how the incorporation of machine learning and uncertainty influences decision rankings. This research advances population distribution modeling by offering a robust, quantitative approach that explicitly accounts for uncertainty in the underlying data, along with guidance for how users can apply uncertainty in their decision-making.

Environmental sciences↗

Improving the LandScan USA Non-Obligate Population Estimate (NOPE)

Where do people go when they have nowhere to be? Nonobligate activities are a significant part of our social and cultural lives, but there are no existing large scale data which characterize spatial variability in population allocation for these activities. As large scale population estimates have ever-finer resolutions, gaps in our ability to estimate this population segment have an increasingly large impact on high resolution population estimates. In this paper, we demonstrate an improved method for estimating the spatial allocation of the non-obligate population - people who are not at work, school, or in another residential institution. This method builds upon on anonymized and aggregate data on visits to public places, allocating the non-obligate population proportionally to worker population while accounting for the estimated ratio of visitors to workers in public places.

Brelsford, Christa↗

Effect of Image Classification Accuracy on Dasymetric Population Estimation

Dasymetric mapping involves the disaggregation of count data, usually relating to population/demographics, from census enumeration areas to smaller target zones with the aid of an ancillary layer related to population density. The ancillary layer is often a binary classification such as developed versus undeveloped, building versus non-building, and residential versus non-residential, in which one class is treated as populated and the other as unpopulated. While dasymetric mapping relies heavily on ancillary data, little research has been done to address the error in ancillary data and its effects on dasymetric mapping accuracy. This chapter reports our research effort to investigate the effect of image classification accuracy on dasymetric population estimates by developing a binomial classification of buildings from high-resolution remote sensor imagery. The classifier was systematically and iteratively manipulated to generate a series of outputs with variegated accuracy characteristics. Lastly, we generated a corresponding series of population estimates based on the mapped building area from each iteration to investigate the relationship between the accuracy of classification and population estimation.

McKee, Jacob↗

At Risk Population Estimates for Belarus, Poland and Slovakia with Machine Learning

High-resolution gridded population modeling is crucial for various applications, including disaster response planning, infectious disease spread modeling, climate change impact estimation, policy development, and more. Multiple gridded population datasets have been developed, each tailored to meet specific objectives. Among them, LandScan Global dataset is designed to represent ambient and unwarned population distributions. However, this dataset relies on a statistical approach that requires manual adjustments, making it time consuming and labour intensive. Existing machine learning (ML) methods often train and test at different spatial resolutions, potentially leading to inflated results, and they rely on Census population totals for disaggregation. To address these limitations, in this study we developed population estimates using ML models trained and tested at a consistent 30 arc-second resolution (≈1 square kilometer), specifically using Random Forest (RF) and XGBoost. These models were trained on 2020 datum to predict for 2021 for three countries: Belarus, Poland, and Slovakia. Our findings show that both RF (MAE varies from 5.75 to 13.25) and XGBoost (MAE varies from 8.15 to 23.44) model performance is close to LandScan Global estimates. Furthermore, neither of the models performed the best across all grid cells: the RF model was more effective in areas with lower populations, while XGBoost excelled in more densely populated regions. The proposed approach can be used for countries where the Census data is not available.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Disturbance of hibernating bats due to researchers entering caves to conduct hibernacula surveys

Estimating population changes of bats is important for their conservation. Population estimates of hibernating bats are often calculated by researchers entering hibernacula to count bats; however, the disturbance caused by these surveys can cause bats to arouse unnaturally, fly, and lose body mass. We conducted 17 hibernacula surveys in 9 caves from 2013 to 2018 and used acoustic detectors to document cave-exiting bats the night following our surveys. We predicted that cave-exiting flights (i.e., bats flying out and then back into caves) of Townsend’s big-eared bats (Corynorhinus townsendii) and western small-footed myotis (Myotis ciliolabrum) would be higher the night following hibernacula surveys than on nights following no surveys. Those two species, however, did not fly out of caves more than predicted the night following 82% of surveys. Nonetheless, the activity of bats flying out of caves following surveys was related to a disturbance factor (i.e., number of researchers × total time in a cave). We produced a parsimonious model for predicting the probability of Townsend’s big-eared bats flying out of caves as a function of disturbance factor and ambient temperature. That model can be used to help biologists plan for the number of researchers, and the length of time those individuals are in a cave to minimize disturbing bats.

59 BASIC BIOLOGICAL SCIENCES↗

National population mapping from sparse survey data: A hierarchical Bayesian modeling framework to account for uncertainty

Population estimates are critical for government services, development projects, and public health campaigns. Such data are typically obtained through a national population and housing census. However, population estimates can quickly become inaccurate in localized areas, particularly where migration or displacement has occurred. Some conflict-affected and resource-poor countries have not conducted a census in over 10 y. We developed a hierarchical Bayesian model to estimate population numbers in small areas based on enumeration data from sample areas and nationwide information about administrative boundaries, building locations, settlement types, and other factors related to population density. We demonstrated this model by estimating population sizes in every 10- m grid cell in Nigeria with national coverage. These gridded population estimates and areal population totals derived from them are accompanied by estimates of uncertainty based on Bayesian posterior probabilities. The model had an overall error rate of 67 people per hectare (mean of absolute residuals) or 43% (using scaled residuals) for predictions in out-of-sample survey areas (approximately 3 ha each), with increased precision expected for aggregated population totals in larger areas. This statistical approach represents a significant step toward estimating populations at high resolution with national coverage in the absence of a complete and recent census, while also providing reliable estimates of uncertainty to support informed decision making.

99 GENERAL AND MISCELLANEOUS↗

TOWARDS RAPID RESPONSE UPDATES OF POPULATIONS AT RISK

Understanding population at risks has been a focus of the LandScan program through its development of population estimates. With advancements in computer vision, deep learning technologies and access to High Performance Computing (HPC) and high resolution imagery, population estimates are now modeled at the building level. However, when those patterns are disrupted, rapid updates to population distribution estimates are needed to support humanitarian aid and response. Oak Ridge National Laboratory (ORNL) recently adapted an existing deep learning building footprint extraction model in development of a scalable approach to Building Damage Assessments (BDA). This new opportunity opens the possibility of automating BDA to support rapid population distribution estimate updates for geographic areas involved in geopolitical conflicts or natural events for humanitarian aid and response or where to focus recovery efforts. In addition, incorporate social surveys to further model human behavior under conflict or other scenarios that disrupt normal patterns of life.

Urban, Marie↗

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

Post-Release Survivorship in a New Population of Blanding's Turtles Established Using Headstarted and Directly Released Turtles

Blanding’s Turtles are facing a variety of anthropogenic threats that decrease population viability. We initiated a project to establish a new population of Blanding’s Turtles on a National Wildlife Refuge in Massachusetts using hatchlings obtained from a nearby robust donor population. We released 440 head-started individuals and 401 directly-released hatchlings between 2007–2013 and conducted multiple years of post-release monitoring via aquatic trapping to estimate survival. Between 2014–2021 we released an additional 824 head-started and 433 directly-released individuals and conducted one year of aquatic trapping in 2021 to estimate population size. First year post-release survival of head-started turtles was six times that of directly-released hatchlings (0.72 vs. 0.12, respectively), but annual survival of both groups was 0.78–0.90 in subsequent years. Only 19% of turtles released at <60 mm carapace length (CL) were subsequently recaptured compared to 60% of those released at ≥60 mm CL, supporting 60 mm CL as a target minimum release size. The mean population estimates using two different study designs were 164 and 202, with 91 individuals encountered. Now that turtles released early in the project are approaching reproductive maturity, we recommend resampling and expanding trapping efforts to include more wetlands to assess long-term survival, abundance, and distribution of Blanding’s Turtles in this newly established population. Furthermore, our results, combined with other evaluations of head-starting across the species’ range, add credence to the use of head-starting as a population recovery tool for Blanding’s Turtles.

59 BASIC BIOLOGICAL SCIENCES↗

Satellite Embedding-Based Population Imputation for Areas with Missing Building Footprint Data: A Computer Vision-Based Approach

High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.

97 MATHEMATICS AND COMPUTING↗

Estimating the population level impact of a gonococcal vaccine candidate: Predictions from a simple mathematical model

Neisseria gonorrhoeae cross-protection was suggested in a New Zealand meningitis B vaccine. We modeled the potential impact of similar vaccines on gonorrhea prevalence in heterosexuals in the United States. Here, our mathematical model incorporated infection, behavior, and vaccination dynamics. Approximate Bayesian Computation calibrated our model to US prevalence. Primary analyses assumed New Zealand vaccine characteristics: 30% efficacy and 2-year duration of protection. We estimated impact under two vaccine coverages (20%, 50%). Reduction in gonorrhea prevalence ranged from 4.8 to 39.4%, depending on vaccine coverage. Vaccine impact was correlated with both size of the highly sexually active subpopulation and sexual mixing between high and low activity subpopulations. A meningitis vaccine providing low efficacy cross-protection against gonorrhea acquisition and short duration of protection could result in a large reduction in gonorrhea prevalence in the United States. Potential dual protective effects can be considered when making vaccine recommendations.

60 APPLIED LIFE SCIENCES↗

Global divergence in urban demographic change and migration patterns

Cities are central to economic development, climate adaptation and social stability, yet globally consistent evidence on how city populations are changing remains limited. Here we analyze annual age- and sex-structured population estimates for more than 10,000 cities worldwide from 2000 to 2020 and show that urban demographic change was highly uneven. Globally, the ratio of children and older adults to working-age adults declined from 0.87 to 0.59, but smaller cities remained consistently younger than larger cities, especially in Africa. We also find pronounced spatial variation in urban sex ratios, including strong male surpluses in parts of the Middle East and North Africa, consistent with patterns of labor migration. Finally, we estimate that 45% of urban population growth was attributable to net migration and 55% to natural increase. These results show that national averages can obscure substantial differences between cities, and highlight the value of globally consistent city-level demographic estimates for understanding regional demographic change and informing locally tailored urban planning.

development studies↗

Sporophyte Stage Genes Exhibit Stronger Selection Than Gametophyte Stage Genes in Haplodiplontic Giant Kelp

Macrocystis pyrifera (giant kelp), a haplodiplontic brown macroalga that alternates between a macroscopic diploid (sporophyte) and a microscopic haploid (gametophyte) phase, provides an ideal system to investigate how ploidy background affects the evolutionary history of a gene. In M. pyrifera , the same genome is subjected to different selective pressures and environments as it alternates between haploid and diploid life stages. We assembled M. pyrifera gene models using available expression data and validated 8,292 genes models using the model alga Ectocarpus siliculosus . Differential expression analysis identified gene models expressed in either or both the haploid and diploid life stages while functional annotation identified processes enriched in each stage. Genes expressed preferentially or exclusively in the gametophyte stage were found to have higher nucleotide diversity (π = 2.3 × 10 –3 and 2.8 × 10 –3 , respectively) than those for sporophytes (π = 1.1 × 10 –3 and 1 × 10 –3 , respectively). While gametophyte-biased genes show faster sequence evolution, the sequence evolution exhibits less signatures of adaptations when compared to sporophyte-biased genes. Our findings contrast the standing masking hypothesis, which predicts higher standing genetic variation at the sporophyte stage, and support the strength of expression theory, which posits that genes expressed more strongly are expected to evolve slower. We argue that the sporophyte stage undergoes more stringent selection compared with the gametophyte stage, which carries a heavy genetic load associated with broadcast spawning. Furthermore, using whole-genome sequencing, we confirm the strong population structure in wild M. pyrifera populations previously established using microsatellite markers, and estimate population genetic parameters, such as pairwise genetic diversity and Tajima’s D , important for conservation and domestication of M. pyrifera .

Molano, Gary↗

Mapathons versus automated feature extraction: a comparative analysis for strengthening immunization microplanning

Background: Social instability and logistical factors like the displacement of vulnerable populations, the difficulty of accessing these populations, and the lack of geographic information for hard-to-reach areas continue to serve as barriers to global essential immunizations (EI). Microplanning, a population-based, healthcare intervention planning method has begun to leverage geographic information system (GIS) technology and geospatial methods to improve the remote identification and mapping of vulnerable populations to ensure inclusion in outreach and immunization services, when feasible. We compare two methods of accomplishing a remote inventory of building locations to assess their accuracy and similarity to currently employed microplan line-lists in the study area. Methods: The outputs of a crowd-sourced digitization effort, or mapathon, were compared to those of a machine-learning algorithm for digitization, referred to as automatic feature extraction (AFE). The following accuracy assessments were employed to determine the performance of each feature generation method: (1) an agreement analysis of the two methods assessed the occurrence of matches across the two outputs, where agreements were labeled as “befriended” and disagreements as “lonely”; (2) true and false positive percentages of each method were calculated in comparison to satellite imagery; (3) counts of features generated from both the mapathon and AFE were statistically compared to the number of features listed in the microplan line-list for the study area; and (4) population estimates for both feature generation method were determined for every structure identified assuming a total of three households per compound, with each household averaging two adults and 5 children. Results: The mapathon and AFE outputs detected 92,713 and 53,150 features, respectively. A higher proportion (30%) of AFE features were befriended compared with befriended mapathon points (28%). The AFE had a higher true positive rate (90.5%) of identifying structures than the mapathon (84.5%). The difference in the average number of features identified per area between the microplan and mapathon points was larger (t = 3.56) than the microplan and AFE (t = -2.09) (alpha = 0.05). Conclusions: Our findings indicate AFE outputs had higher agreement (i.e., befriended), slightly higher likelihood of correctly identifying a structure, and were more similar to the local microplan line-lists than the mapathon outputs. These findings suggest AFE may be more accurate for identifying structures in high-resolution satellite imagery than mapathons. However, they both had their advantages and the ideal method would utilize both methods in tandem.

59 BASIC BIOLOGICAL SCIENCES↗

LandScan HD: a high-resolution gridded ambient population methodology for the world

Unwarned population distributions accounting for routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (LSHD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (approximately 90 m). Although LSHD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LSHD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LSHD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LSHD baseline dataset, we contrast unwarned and residential (WorldPop) population distributions by (1) exploring a practical application of flood risk assessment and (2) evaluating their congruence with outcomes of collective human activities (subnational CO 2 emissions). Finally, we discuss plans to address current LSHD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.

Building morphology↗

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics↗