Engineering PapersSearch

SEARCH · Engineering Papers

Results for “population forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PopGNN: Graph Neural Network-Based Flexible Future Population Forecasting Model

Accurate population forecasts is important to plan critical infrastructure and services, from housing and education to healthcare and transport. However, traditional population prediction studies have only employed traditional machine learning models limited to capture complex spatial interdependencies and patterns. Althogh recently computer vision-based framework was introduced with with promising accuracy, it has critical limitations for real-world planning applications: it function only at fixed spatial resolutions, restricting their use in diverse boundaries such as census tracts, neighborhoods, or administrative zones. Therefore, this study suggests a Graph Neural Network (GNN)-based population prediction framework, called PopGNN. This model recorded remarkable performance compared with state-of-the-art models and traditional baseline models in the grid and administrative boundaries. Furthermore, our framework achieved comparable predictive accuracy to a computer vision-based model in both the South Korea and Tennessee case studies. Consequently, this study is valuable in that a single model can provide accurate population forecasts that address diverse planning demands, ranging from granular grid-level estimates for precise service allocation and facility location planning to aggregate administrative-level forecasts for macro-scale regional policy and resource distribution.

97 MATHEMATICS AND COMPUTING

Popnet : computer vision based deep learning model for forecasting gridded population

Here, this study introduces Popnet, a deep learning model for forecasting 1 km-gridded populations, integrating U-Net, ConvLSTM, a Spatial Autocorrelation module and deep ensemble methods. Using spatial variables and population data from 2000 to 2020, Popnet predicts South Korea’s population trends by age groups (under 14, 15-64 and over 65) up to 2040. In validation, it outperforms traditional machine learning and state-of-the-art computer vision models. The output of this model discovered significant polarisation: population growth in urban areas, especially the capital region, and severe depopulation in rural areas. Popnet is a robust tool for offering significant insights to policymakers and related stakeholders about the detailed future population, which allows them to establish detailed, localised planning and resource allocations.

computer vision

Evaluation of the 2022 West Nile virus forecasting challenge, USA

Abstract Background West Nile virus (WNV) is the most common cause of mosquito-borne disease in the continental USA, with an average of ~1200 severe, neuroinvasive cases reported annually from 2005 to 2021 (range 386–2873). Despite this burden, efforts to forecast WNV disease to inform public health measures to reduce disease incidence have had limited success. Here, we analyze forecasts submitted to the 2022 WNV Forecasting Challenge, a follow-up to the 2020 WNV Forecasting Challenge. Methods Forecasting teams submitted probabilistic forecasts of annual West Nile virus neuroinvasive disease (WNND) cases for each county in the continental USA for the 2022 WNV season. We assessed the skill of team-specific forecasts, baseline forecasts, and an ensemble created from team-specific forecasts. We then characterized the impact of model characteristics and county-specific contextual factors (e.g., population) on forecast skill. Results Ensemble forecasts for 2022 anticipated a season at or below median long-term WNND incidence for nearly all (> 99%) counties. More counties reported higher case numbers than anticipated by the ensemble forecast median, but national caseload (826) was well below the 10-year median (1386). Forecast skill was highest for the ensemble forecast, though the historical negative binomial baseline model and several team-submitted forecasts had similar forecast skill. Forecasts utilizing regression-based frameworks tended to have more skill than those that did not and models using climate, mosquito surveillance, demographic, or avian data had less skill than those that did not, potentially due to overfitting. County-contextual analysis showed strong relationships with the number of years that WNND had been reported and permutation entropy (historical variability). Evaluations based on weighted interval score and logarithmic scoring metrics produced similar results. Conclusions The relative success of the ensemble forecast, the best forecast for 2022, suggests potential gains in community ability to forecast WNV, an improvement from the 2020 Challenge. Similar to the previous challenge, however, our results indicate that skill was still limited with general underprediction despite a relative low incidence year. Potential opportunities for improvement include refining mechanistic approaches, integrating additional data sources, and considering different approaches for areas with and without previous cases. Graphical Abstract

54 ENVIRONMENTAL SCIENCES

CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting

With hundreds of SARS-CoV-2 lineages circulating in the global population, there is an ongoing need for predicting and forecasting lineage frequencies and thus identifying rapidly expanding lineages. Accurate prediction would allow for more focused experimental efforts to understand pathogenicity of future dominating lineages and characterize the extent of their immune escape. Here, we first show that the inherent noise and biases in lineage frequency data make a commonly-used regression-based approach unreliable. To address this weakness, we constructed a machine learning model for SARS-CoV-2 lineage frequency forecasting, called CovTransformer, based on the transformer architecture. We designed our model to navigate challenges such as a limited amount of data with high levels of noise and bias. We first trained and tested the model using data from the UK and the USA, and then tested the generalization ability of the model to many other countries and US states. Remarkably, the trained model makes accurate predictions two months into the future with high levels of accuracy both globally (in 31 countries with high levels of sequencing effort) and at the US-state level. Our model performed substantially better than a widely used forecasting tool, the multinomial regression model implemented in Nextstrain, demonstrating its utility in SARS-CoV-2 monitoring. Assuming a newly emerged lineage is identified and assigned, our test using retrospective data shows that our model is able to identify the dominating lineages 7 weeks in advance on average before they became dominant. Overall, our work demonstrates that transformer models represent a promising approach for SARS-CoV-2 forecasting and pandemic monitoring.

60 APPLIED LIFE SCIENCES

Propagating synthetic populations with dynamic Bayesian networks: a framework for long-horizon demographic forecasting

This study presents a dynamic demographic microsimulator using dynamic Bayesian networks to forecast long–term changes in household and individual life events. Leveraging longitudinal Panel Study of Income Dynamics (PSID) data, two networks for individuals and households were modeled to simulate transitions in employment, income, education, marriage, childbirth, leaving the parental home, home ownership, mortality, and household formation or dissolution. Across 1,000 simulation runs spanning 24 years, household–level outcomes remain highly accurate and individual–level predictions reasonable. Although accuracy naturally declines with projection horizon, performance remains promising at both levels. This study addresses a key limitation of existing population synthesis models, which typically generate only a single static snapshot of the population. In conclusion, by introducing a framework that propagates cross-sectional outputs into the future, the microsimulator enables the tracking of demographic evolution over time, enhances realism in population-based simulations, and supplies credible inputs to agent-based travel demand models.

Demographic modeling

An integrated integral projection model ( IPM 2 ) to disentangle size‐structured harvest and natural mortality

Abstract Body size is one of the most important traits governing individual‐level demographic rates and modulating population‐level processes. Multiple size‐dependent demographic rates can simultaneously change population structure, so distinguishing their individual contributions to overall population dynamics remains a challenge. Disentangling size‐dependent harvest rates from other demographic rates is critical for assessing the impact of removal on populations of invasive species. Inference about invasive populations can be difficult, however, as observations are often collected opportunistically as part of removal programs, rather than experimentally designed. Yet accurate inference is essential for understanding the feasibility of population suppression and optimising management decisions. We develop an integrated integral projection model (IPM 2 ) that leverages the strengths of the integrated population model and integral projection model to enable inference about complex, size‐structured demographic rates from imperfect observations. We apply the IPM 2 in the context of invasive European green crab ( Carcinus maenas ), a species for which individual body size strongly regulates both the observation‐generating process and latent, population dynamics. The IPM 2 facilitates the distinct estimation of green crab size‐structured harvest and natural mortality rates, parameters for which no explicit data is collected and that are unidentifiable in component datasets of the integrated population model. The model represents how the green crab population changes over time, providing the first estimates of size‐structured abundance of this high‐priority species. By forecasting the stable size distribution and equilibrium population size under varying removal efforts, we demonstrate that extremely high levels of removal effort can reduce the equilibrium green crab population size. Yet these high mortality rates also shift the stable size distribution and increase the equilibrium abundance of smaller crabs, since size‐selective removal alters intraspecific interactions. The ecological outcome of this shift in size structure will be variable, as green crab size modulates only some of its interactions with other species. These results highlight the value of the IPM 2 framework for inferring complex population dynamics with information needs that outpace information in individual observational datasets, providing a path forward for accurate assessment of conservation programs.

Keller, Abigail G. [Department of Environment Scie

Gravitational waves and galaxies cross-correlations: a forecast on GW biases for future detectors

ABSTRACT Gravitational waves (GWs) have rapidly become important cosmological probes since their first detection in 2015. As the number of detected events continues to rise, upcoming instruments like Einstein Telescope (ET) and Cosmic Explorer (CE) will observe millions of compact binary mergers. These detections, coupled with galaxy surveys by instruments such as the Dark Spectroscopic Energy Instrument (DESI), Euclid, and the Vera Rubin Observatory, will provide unique information on the large-scale structure of the universe by cross-correlating GWs with the distribution of galaxies hosting them. In this paper, we focus on how cross-correlations constrain the clustering bias of GWs emitted by the coalescence of binary black holes (BBHs). This parameter links BBHs to the underlying dark matter distribution, hence informing us how they populate galaxies. Using a multitracer approach, we forecast the precision of these measurements under different survey combinations. Our results indicate that current GW detectors will have limited precision, with measurement errors as high as $\displaystyle \sim 50~{{\ \rm per\ cent}}$. However, third-generation detectors like ET, when cross-correlated with Legacy Survey of Space and Time (LSST) data, can improve clustering bias measurements to within 2.5 per cent. Furthermore, we demonstrate that these cross-correlations can enable a per cent-level measurement of the magnification lensing effect on GWs. Despite this, there is a degeneracy between magnification and evolution biases, which hinders the precision of both. This degeneracy is most effectively addressed by assuming knowledge of one bias or targeting an optimal redshift range of $\displaystyle 1 \lt z \lt 2.5$. Our analysis opens new avenues for studying the distribution of BBHs and testing the nature of gravity through large-scale structure.

Zazzera, Stefano (ORCID:0000000158979221)

EV Load Forecasting Guide: A Report by the Energy Systems Integration Group’s EV Load Forecasting Task Force

Forecasting electricity usage is a foundational planning activity for utilities, underpinning billions of dollars in grid investments that ensure system reliability. Historically, forecasting relied on trends in economic and population growth; however, transportation electrification presents a new and complex planning challenge. Unlike conventional loads, electric vehicle (EV) charging has relatively limited usage history. In addition, charging is driven by complex human behaviors, is mobile, and at the same time can concentrate geographically in ways that, without proper planning, can quickly overwhelm local distribution systems.

Giraldez, Julieta [Electric Power Engineers, Austi

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University

Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases

A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.

97 MATHEMATICS AND COMPUTING

Multiscale drivers of extreme southern California flooding: ENSO, MJO, North Pacific jet, and atmospheric rivers

Extreme rainfall and flooding, driven by a powerful atmospheric river (AR) and a persistent Madden-Julian Oscillation (MJO), hit Southern California in February 2024 during the 2023–2024 El Niño, affecting over 10 million people. ARs are key contributors to extreme rainfall and flooding along the U.S. West Coast. Although the AR-MJO link has been documented, its spatio-temporal variability remains a major forecasting and risk-management challenge. Combining precipitation, stream gauge and demographic data, we quantify the physical drivers and population exposure to this extreme event. Leveraging a Lagrangian MJO precipitation tracking algorithm, we unravel the multiscale interactions responsible for the AR’s development. El Niño favored a large, long-lived MJO that interacted with the North Pacific Jet (NPJ) over more than three weeks. The MJO convective outflow modulated the NPJ by inducing negative potential vorticity advection along the tropopause. The ensuing NPJ extension and acceleration induced explosive cyclogenesis, whose AR-driven moisture transport resulted in extreme rainfall.

Atmospheric dynamics

Optimizing resource allocation in Miscanthus breeding via sparse testing designs for genomic prediction

Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in conventional perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse bio-based products. Increasing biomass yield will increase profitability and environmental benefits, so it is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed implementing sparse testing designs. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction consistently presented the highest PA and the lowest MSE: for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.

Miscanthus sacchariflorus (MSA)

Neutrino and gamma-ray signatures of inelastic dark matter annihilating outside neutron stars

We present a new inelastic dark matter search: neutron stars in dark matter-rich environments capture inelastic dark matter which, for interstate mass splittings between about 45 - 285 MeV, will annihilate away before becoming fully trapped inside the object. This means a sizable fraction of the dark matter particles can annihilate while being outsidethe neutron star, producing neutron star-focused gamma-rays and neutrinos. We analyze this effect for the first time and target the neutron star population in the Galactic Center, where the large dark matter and neutron star content makes this signal most significant. Depending on the assumed neutron star and dark matter distributions, we set constraints on the dark matter-nucleon inelastic cross-section using existing H.E.S.S. observations. We also forecast the sensitivity of upcoming gamma-ray and neutrino telescopes to this signal, which can reach inelastic cross-sections as low as ∼ 2 × 10 -47 cm 2 .

79 ASTRONOMY AND ASTROPHYSICS

Analyzing Trip Chaining Behavior in New York State Using 2009 and 2017 National Household Travel Survey

Trip chaining, defined as the sequential linking of trips by individuals throughout a given day, provides critical insights into daily mobility patterns and activity sequencing. Understanding these patterns has significant implications for transportation demand forecasting, congestion management, and local economic activity. This analysis examines trip chaining behaviors in New York State (NYS) for the years 2009 and 2017 and compares the Middle Atlantic Census Division with other U.S. regions in 2022, utilizing data from the National Household Travel Survey (NHTS). Through demographic, geographic, and temporal analysis, this study characterizes how populations organize travel for work, personal errands, and social activities, providing empirical evidence of evolving trip chaining behaviors to inform transportation planning strategies.

99 GENERAL AND MISCELLANEOUS

Graph-Based Prediction of Spatio-Temporal Vaccine Hesitancy From Insurance Claims Data

Growing vaccine hesitancy is contributing to the decline in immunization rates for highly contagious, vaccine-preventable childhood diseases. Therefore, there has been a significant interest in understanding how hesitancy is spreading at higher spatio-temporal resolutions, enabling more targeted interventions. Motivated by this, we study the problem of prediction of vaccine hesitancy at the ZIP Code level, referred to as the VaxHesitancy problem. A significant challenge for this problem is the lack of high-resolution data that indicates hesitancy. Here, we develop a hybrid VaxHesSTL framework that combines a Graph Neural Network (GNN) and a Recurrent Neural Network (RNN) to address the VaxHesitancy problem. The GNN uses a ZIP Code-level network to capture spatial signals from neighboring areas, while the RNN models the temporal dynamics present in the data. We train and evaluate VaxHesSTL using a large dataset, namely the All-Payer Claims Databases (APCD), for Virginia, consisting of insurance claims from over five million individuals for six years. We find that an aggregated contact network or graph, developed from a detailed activity-based population network, plays an important role in the performance of VaxHesSTL, compared to graph models based solely on spatial proximity. Experiments demonstrate that VaxHesSTL outperforms a range of state-of-the-art baselines, which rely solely on historical time series data without accounting for spatial relationships. Since hesitancy data at higher spatial resolution is often unavailable or hard to get, we incorporate an active learning approach with our VaxHesSTL framework to optimize the training set without compromising the prediction performance. We find that hesitancy data for only 18% of ZIP Codes selected by active learning allows us to forecast hesitancy for all the ZIP Codes in the Virginia.

60 APPLIED LIFE SCIENCES

Unraveling TeV halos with the Cherenkov Telescope Array

Pulsars are observed to emit bright and spatially extended gamma-ray emission at multi-TeV energies. These so-called "TeV halos" are now understood to be a nearly universal feature of middle-aged pulsars. However, many of the key physical processes that govern these systems, particularly those affecting particle diffusion, remain poorly constrained. We aim to evaluate the ability of the Cherenkov Telescope Array (CTA) to probe the physical properties of TeV halos, with a focus on the nearby and well-studied case of the Geminga pulsar. We simulate gamma-ray emission from various TeV halo models, incorporating different assumptions for the injected electron spectrum, spin-down evolution, and energy-dependent diffusion. These models are then used to forecast CTA's sensitivity to spectral and spatial differences, based on realistic mock observations and instrument response simulations. We find that CTA will be able to distinguish between a wide range of TeV halo models that are currently consistent with existing data. In particular, CTA observations can constrain the normalization, energy dependence, and spatial extent of the diffusion coefficient surrounding Geminga, as well as the spectral shape of the injected electron population.

79 ASTRONOMY AND ASTROPHYSICS

Uncertainties in the effects of organic aerosol coatings on polycyclic aromatic hydrocarbon concentrations and their estimated health effects

We used the CAM5 model to examine how different particle-bound polycyclic aromatic hydrocarbon (PAH) degradation approaches affect the spatial distribution of benzo(a)pyrene (BaP). Three approaches were evaluated: NOA (no effect of OA coatings state on BaP), shielded (viscous OA coatings shield BaP from oxidation under cool and dry conditions) and ROI-T (viscous OA coatings slow BaP oxidation in response to temperature and humidity). Results show that BaP concentrations vary seasonally, influenced by emissions, deposition, transport and degradation approach, all of which are influenced by meteorological conditions. All simulations predict higher population-weighted global average (PWGA) fresh BaP concentrations during December–January–February (DJF) compared to June–July–August (JJA), due to increased emissions from household activities and reduced removal processes during colder months. The shielded and ROI-T approaches, which account for OA coatings, result in 2–6 times higher BaP concentrations in DJF compared to NOA. The shielded simulation predicts the highest PWGA fresh BaP concentration (1.3 ng m −3 ), with 90 % of BaP protected from oxidation. In contrast, the ROI-T approach forecasts lower concentrations in middle to low latitudes, as it assumes less effective OA coatings under warmer, more humid conditions. Evaluations against observed BaP concentrations show the shielded approach performs best, with a normalized mean bias (NMB) within ± 20 %. The combined incremental lifetime cancer risk (ILCR) for both fresh and oxidized PAHs is similar across simulations, emphasizing the importance of considering both forms in health risk assessments. This study highlights the critical role of accurate degradation approaches in PAH modeling.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN