Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “time-series modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Spatiotemporal Learning in Power Modules: Wavelet-Enhanced Forecasting of Thermomechanical Degradation

Detecting internal defects in power electronics packages is critical for their performance and reliability, especially under extreme operating conditions, as these defects can lead to catastrophic failure if not properly addressed. Confocal scanning acoustic microscopy (C-SAM) plays a key role in the nondestructive evaluation of bond layer degradation within a power electronics package by detecting defects such as delamination, voids, and cracks. However, accurately quantifying and predicting these defects from C-SAM images remains a significant challenge due to the low noise-to-signal ratio, which typically arises from both imaging process and bond patterns itself. In this paper, we explore machine learning strategies for processing C-SAM images and providing predictive models of defect growth. We use C-SAM images of sintered copper and sintered silver samples, which are obtained under accelerated thermal experiments, as the representative dataset for our study. We investigate the effect of Fourier transforms and wavelet transforms on these datasets to remove high-frequency noise and address noise across multiple scales with histogram equalization to enhance the contrast and improve the visibility of defects. As a result, defect boundaries can be clearly distinguished, enabling more accurate tracking of their growth over time. We then employ different time-series forecasting algorithms on the denoised images to formulate an image-based lifetime prediction model. Statistical models and deep-learning techniques are trained on images obtained in the early stages of thermal shock, and defect growth in the later stages is predicted. Our work serves as a preliminary attempt to improve the accuracy of lifetime prediction models of power electronics packages, which is critical under extreme operating environments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

sdt (Solar Data Tools) [SWR-25-130]

Solar Data Tools (sdt) is an open-source Python library for analyzing PV power (and irradiance) time-series data. It was developed to enable analysis of unlabeled PV data, i.e. with no model, no meteorological data, and no performance index required, by taking a statistical signal processing approach in the algorithms used in the package’s main data processing pipeline. Solar Data Tools empowers PV system fleet owners or operators to analyze system performance a hundred times faster even when they only have access to the most basic data stream—power output of the system.

Meyers-Im, Bennet [National Laboratory of the Rock↗

Spatiotemporal 4D Whole-cell Modeling of a Minimal Autotroph Reveals Central Carbon Metabolism Regulated Locally by Protein Megacomplexes via Post-translational Modifications under Light Disturbance

Photosynthetic microorganisms rely on multiple pathways in central carbon metabolism to adapt to fluctuating light and energy availability across diel cycles. Mechanistic insight into the regulatory dynamics of this adaptation requires integrating processes spanning disparate timescales, from rapid redox-dependent post-translational modifications (PTMs) to slower changes in protein expression and metabolic pathway usage. To address this complexity beyond genome-based inference and traditional modeling, we develop a whole-cell four-dimensional (3D + time) model of the marine cyanobacterium Prochlorococcus marinus MED4 that explicitly represents the spatial organization of enzymatic and molecular processes in central carbon metabolism under light perturbation. We employ a perturbation-based research design to experimentally generate time-series, multi-omics measurements that provide molecular descriptors and cryo-ET derived 3D segmented volumes as constraints for this dynamic 4D framework. The integration of experiments and modeling across defined light regimes enables quantitative validation of system-level responses and forecasting under distinct light disturbances. We test the hypothesis that light-dependent redox PTMs regulating the structural assembly of a protein megacomplex, the “dark complex,” modulate metabolic flux at a conserved regulatory node of the Calvin–Benson cycle (CBC) in cyanobacteria. Our model shows that subcellular spatial organization buffers rapid light-induced changes in thylakoid reaction rates, which are followed by redox-PTM-mediated sequestration or release of CBC enzymes in the dark complex, ultimately impacting carbon fixation dynamics within carboxysomes. Comparison with an equivalently parameterized well-mixed stochastic model demonstrates that post-translational regulation not only buffers transcriptional noise and diffusion-driven fluctuations but also stabilizes phenotypic outcomes, underscoring the importance of spatial heterogeneity in phenotypic robustness. This ability to probe adaptive, spatiotemporally resolved mechanisms in photosynthetic machinery and central carbon metabolism addresses a critical gap in genotype-to-phenotype inference and expands modeling and design capabilities for understudied or genetically intractable autotrophs such as P. marinus MED4.

Johnson, Connah G.↗

Machine learning pipeline for denoising low signal-to-noise ratio and out-of-distribution transmission electron microscopy datasets

High-resolution transmission electron microscopy (HRTEM) is crucial for observing material’s structural and morphological evolution at Angstrom scales, but the electron beam can alter these processes. Devices such as CMOS-based direct-electron detectors operating in electron-counting mode can be utilized to substantially reduce the electron dosage. However, the resulting images often lead to a low signal-to-noise ratio, which requires frame integration that sacrifices temporal resolution. Several machine learning (ML) models have been recently developed to successfully denoise HRTEM images. Yet, these models are often computationally expensive, and their inference speeds on GPUs are outpaced by the imaging speed of advanced detectors, precluding in situ analysis. Furthermore, the performance of these denoising models on datasets with imaging conditions that deviate from the training datasets has not been evaluated. To mitigate these gaps, we propose a new self-supervised ML denoising pipeline specifically designed for time-series HRTEM images. This pipeline integrates a blind-spot convolution neural network with pre-processing and post-processing steps, including drift correction and low-pass filtering. Results demonstrate that our model outperforms various other ML and non-ML denoising methods in noise reduction and contrast enhancement, leading to improved visual clarity of atomic features. Additionally, the model is drastically faster than U-Net-based ML models and demonstrates excellent out-of-distribution generalization. The model’s computational inference speed is in the order of milliseconds per image, rendering it suitable for application in in-situ HRTEM experiments.

36 MATERIALS SCIENCE↗

Data-driven particle dynamics: Structure-preserving coarse-graining for emergent behavior in non-equilibrium systems

Multiscale systems are ubiquitous in science and technology, but are notoriously challenging to simulate as short spatiotemporal scales must be appropriately linked to emergent bulk physics. When expensive high-dimensional dynamical systems are coarse-grained into low-dimensional models, the entropic loss of information leads to emergent physics which are dissipative, history-dependent, and stochastic. To machine learn coarse-grained dynamics from time-series observations of particle trajectories, we propose a framework using the metriplectic bracket formalism that preserves these properties by construction; most notably, the framework guarantees discrete notions of the first and second laws of thermodynamics, conservation of momentum, and a discrete fluctuation-dissipation balance crucial for capturing non-equilibrium statistics. We introduce the mathematical framework abstractly before specializing to a particle discretization. As labels are generally unavailable for entropic state variables, we introduce a novel self-supervised learning strategy to identify emergent structural variables. We validate the method on benchmark systems and demonstrate its utility on two challenging examples: (1) coarse-graining star polymers at challenging levels of coarse-graining while preserving non-equilibrium statistics, and (2) learning models from high-speed video of colloidal suspensions that capture coupling between local rearrangement events and emergent stochastic dynamics. We provide open-source implementations in both PyTorch and LAMMPS, enabling large-scale inference and extensibility to diverse particle-based systems.

Computational Engineering, Finance, and Science (c↗

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Measurement-informed Dynamic Aggregation of Distribution Systems

This paper proposes a measurement-informed dynamic aggregation methodology in order to create equivalent representations of distribution systems that are compatible with large-scale transmission analysis. By optimizing an equivalent feeder parameters using time-series measurements of active power, reactive power, and voltage at the Point of Interconnection (POI), the approach yields simplified yet dynamically accurate equivalents. Implemented in PSCAD with models of photovoltaic–battery systems, three-phase motors, and static loads, the method employs hybrid differential evolution and bounded least-squares optimization laying the foundation for for real-time state estimation and optimized sensor placement in distribution networks.

Ahmed, Kazi Ishrak [University of Tennessee, Knoxv↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER) With Envelope Improvements and Advanced Air Sealing

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson, et al. 2022). This document focuses on a single end-use savings shape measure: Residential Two-Stage Geothermal Heat Pump (GHP) (4.0 COP, 20.5 EER) With Envelope Improvements. This measure combines a two-stage GHP with envelope improvements as a single package. As this package is a combination of two other measures, this document focused on documenting the results associated with this combination of technologies, with individual measure documents for two-stage GHPs and envelope improvements providing the information on the details of these measures. When the two technologies are combined, envelope improvements can modestly reduce energy consumption by a further 10%-15%, but also reduce the required size of the ground heat exchanger and heat pump by approximately 33% on average across all sites. The cost of installing envelope improvements in these homes is likely to be more than paid for by the reduction in equipment and drilling costs in these buildings for the majority of the stock.

15 GEOTHERMAL ENERGY↗

ResStock Measure Documentation: Residential Single-Stage Geothermal Heat Pump (3.8 COP, 18.6 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This documentation focuses on a single end-use savings shape measure: Residential Single-Stage Geothermal Heat Pump (GHP).?Single-stage GHPs are able to reduce energy consumption by 31% for the entire stock. Additional results provided below detail how savings changes for sections of the housing stock with different base heating fuel and in different climate zones, as well as the savings potential by state for both heating and cooling. Utility bills and electric panel impacts are also shown and discussed.

15 GEOTHERMAL ENERGY↗

ResStock Measure Documentation: Residential Variable-Speed Geothermal Heat Pump (4.4 COP, 30.9 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock (TM) is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This documentation focuses on a single end-use savings shape measure: Residential Variable-Speed Geothermal Heat Pump (GHP). This document provides the relevant new modeling information for variable-speed systems not previously covered in either the single-stage or two-stage documents. Variable-speed GHPs represent the most efficient option available for this technology: They provide the most savings, with up to 46% for the applicable portion of the housing stock, compared to 31% for less efficient single-stage GHPs. Additional results shown here detail how the savings change for sections of the housing stock with different base heating fuels and in different climate zones, and they show the savings potential by state for both heating and cooling. Utility bills and electric panel impacts are also shown and discussed.

15 GEOTHERMAL ENERGY↗

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This document focuses on a single end-use savings shape measure: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER). This document builds on details established in the single-stage document (Maguire et al. 2025) to detail differences in the approach to modeling this higher efficiency, but more commonly deployed, type of geothermal heat pump. Specific EnergyPlus objects and product specific curves used are highlighted along with showing the results of this measure compared to the baseline and single-speed geothermal heat pumps. Two-speed geothermal heat pumps are able to save even more energy and on utility bills than single-speed products, albeit at the expense of a higher first cost.

15 GEOTHERMAL ENERGY↗

Poster Abstract: Leveraging Large Language Models to Reveal Interpretable Cooling Behaviors from Smart Thermostat Data

Frequent heatwaves and hot summers increasingly challenge occupant comfort, health, and energy grid stability. Addressing these challenges requires a detailed understanding of household cooling behaviors, such as thermostat adjustments and adaptive responses to extreme conditions. Traditional analyses often rely on aggregated numerical metrics that overlook subtle but important household-specific variations. In this study, we introduce a generalizable methodology that integrates large language models (LLMs) with vision capabilities to enable scalable and detailed analysis of residential thermostat data. Using Ecobee's Donate Your Data (DYD) dataset—which provides five-minute records of indoor temperatures, thermostat setpoints, and HVAC runtimes—we focus on two U.S. cities with contrasting summer climates : Austin (TX) and Phoenix (AZ). Because raw time-series data are not well suited for direct LLM analysis, we transform them into visual representations, such as daily indoor temperature trajectories and weekly runtime histograms, to better capture behavioral variations. Leveraging LLMs' visual interpretation, we extract descriptive behavioral features, including temperature preferences, time-of-day cooling orientation, anticipatory versus reactive heatwave responses, and behavioral consistency. These semantic features support unsupervised clustering to identify distinct occupant archetypes at scale, revealing differences—such as morning-centric anticipatory coolers versus households that shift toward warmer setpoints during heatwaves—that can inform demand response, resilience planning, and health-aware interventions. By converting raw numerical data into interpretable behavioral patterns, this methodology enables scalable and practical analysis of occupant behavior, supporting actionable insights for comfort, resilience, and energy management.

Nihar, Kopal↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Neural Optimizer With Decision-Focused Learning for Optimal Energy Storage Operation

Here, this article introduces a neural optimizer-based framework for optimizing battery energy storage system (BESS) control for grid services, including demand charge and energy cost reduction. By leveraging decision-focused learning (DFL), the proposed framework ensures seamless integration and adaptation, significantly enhancing control performance. A patch time-series transformer is employed for peak load forecasting, incorporating aleatoric uncertainty quantification to account for forecasting uncertainties within the decision-making process. The framework utilizes a solver-in-the-loop approach to generate optimal BESS actions, which are then used to train the neural optimizer-based agent. By co-optimizing both BESS operational modes and output power within the NN, the system achieves improved performance and robustness. After initial training, the forecasting and control models are jointly fine-tuned to account for forecasting errors, further improving decision precision and efficiency through DFL. Case studies are performed to validate the performance of the framework using multiple real-world datasets, demonstrating superior performance in monthly peak load forecasting compared to state-of-the-art models. In addition, the results are compared against existing decision-making approaches. The results demonstrate a reduction in monthly peak forecasting error by approximately 15% across various performance measures and achieve an optimization gap for BESS operation that is about three times smaller compared to existing methods.

Kim, Hyeonjin [Pacific Northwest National Laborato↗

ComStock Measure Scenario Documentation: Standard Performance Heat Pump Rooftop Unit With New Windows

Building on a 3-year effort to calibrate and validate the U.S. Department of Energy's ResStock (TM) and ComStock (TM) models, this work produces national datasets that empower analysts working for federal, state, utility, city, and manufacturer stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual energy consumption (at subhourly resolution) of the commercial building stock across the United States. The baseline model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology and results of the baseline model are discussed in the final technical report of the End-Use Load Profiles project. The goal of this work is to develop energy efficiency and demand flexibility end-use load shapes that cover high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to various "what-if" scenarios that can be applied to buildings. An end-use savings shape is the difference in energy consumption between a baseline building (or collection of buildings) and a building with an energy efficiency or demand flexibility measure applied. It results in a time-series profile broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step, as well as annual aggregations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University ↗