Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

When to vaccinate for seasonal influenza: check the peak forecast

Background Seasonal influenza infects 5-20% of people every year in the United States, resulting in hospitalizations, deaths, and adverse economic impacts. To mitigate these impacts, influenza vaccines are developed and distributed annually; however, growing evidence suggests that vaccine effectiveness (VE) wanes over the course of a flu season. Delaying influenza vaccination for older adults has attracted attention as a potential public health strategy. However, given the uncertainties in seasonal peak, vaccine effectiveness, and waning rates, postponing vaccination could also lead to increased morbidity, motivating an evaluation of a range of potential scenarios. The aim of this study was to investigate favorable age group-specific vaccination schedules that could lead to the greatest disease burden reduction. Methods We systematically investigated a broad range of vaccination start times for five age groups under six combinations of initial effectiveness and waning rates, based on influenza cases and vaccine uptake data from 10 influenza seasons. We defined the most favorable vaccination schedule as the one that resulted in the greatest reduction in disease burden. Results In scenarios with fast waning, all age groups benefit from delaying vaccination regardless of initial VE and peak timing. In scenarios with slower waning, results are mixed. For the ≥65 group, high initial VE and slow waning suggests that in early-peaking seasons, early vaccination most effectively reduces disease burden, while in late-peaking seasons delaying vaccination is most effective. For the ≥65 group in medium and low initial VE, and slow waning scenarios, delaying vaccination appears to prevent the greatest number of cases, regardless of whether the season peaks early or late. Conclusion The most favorable vaccination schedule is sensitive to changes in initial VE, waning rate, and peak timing. Given estimates of these quantities from statistical and immunological models and observations, our methods can inform vaccination recommendations in order to most effectively reduce the annual disease burden caused by seasonal influenza. Specifically, accurate peak timing forecasts for the upcoming season have the potential to guide decisions on when to vaccinate.

59 BASIC BIOLOGICAL SCIENCES↗

Can Simple Metrics Identify the Process(es) Driving Extreme Precipitation?

This work seeks an automatic algorithm to determine the primary meteorological cause(s) of individual extreme precipitation events. Such determinations have been made before, but required a by-hand analysis of each separate event. This is very time-consuming and the field would benefit from an automatic process. This is especially relevant when comparing different datasets to determine which ones most closely hew towards reality. This paper tests three simple metrics over the continental United States using the European Center for Medium-Range Weather Forecasting’s (ECMWF) atmospheric reanalysis (ERA5). The metrics tested measure and compare the strength of three meteorological processes associated with extreme precipitation: fronts, convection, and cyclones. A multivariate statistical technique as well as individual case studies show evidence that the three meteorological processes of interest cannot be isolated from one another using these simple physical metrics. This shows the difficulty in finding “pure” cases of these precipitation-generating processes and suggests approaching these processes with an eye toward mixed-type events.

Swenson, Leif M. (ORCID:0000000199708735)↗

Physics-based Induced Earthquake Forecasting: Process Understanding, and Hazards Mitigation

Disposal of saltwater co-produced with oil and gas is linked to elevated seismicity in the Central and Midwest US. There is a concern that these events may lead to widespread damage and an overall increase in seismicity. Thus an improved understanding of the spatially and temporally variable deformation and stress field associated with fluid injection operation is critically necessary for evaluating time-varying seismic hazards. Despite the improvements in seismic monitoring capacity and the resulting decrease in the magnitude detection threshold, estimates of induced earthquake probability remain elusive due to insufficient models incapable of accounting for the complex physics governing the process of induced seismicity. The proposed research effort will comprehensively analyze, integrate, and interpret geodetic, injection, and seismic data in the vicinity of the injection sites in Oklahoma to resolve the 4-dimensional distribution of pore pressure and stress in the shallow crust. This project, in particular, is focused on exploring the statistical relation between injection operation and increased earthquake hazard. The amplitude of and the extent to which pore pressure changes are determined by some factors, in particular, the hydrogeological properties of the rocks, such as diffusivity. Thus the available deformation data will be used to constrain hydrogeological properties of the medium, to accurately resolve the evolution of crustal stresses due to fluid injection. Having the time-varying models of stress changes, a statistical framework will be implemented to estimate the time-dependent probability of large earthquakes on the nearby fault systems. These data and models help to improve seismic hazard estimates and aid in constructing operational-induced earthquake forecast models. This information can also be integrated into the updated U.S. National Seismic Hazard Map, which local communities and authorities use in their earthquake risk estimates and mitigation efforts.

58 GEOSCIENCES↗

Large‐Scale Statistically Meaningful Patterns (LSMPs) Associated With Precipitation Extremes Over Northern California

Abstract We analyze large‐scale statistically meaningful patterns (LSMPs) that precede extreme precipitation (PEx) events over Northern California (NorCal). We find LSMPs by applying k‐means clustering to the two leading principal components of daily 500 hPa geopotential height anomalies two days before the onset, from October to March during 1948–2015. Statistical significance testing based on Monte Carlo simulations suggests a minimum of four statistically distinguished LSMP clusters. The four LSMP clusters are characterized as Northwest continental negative height anomaly, Eastward positive “Pacific‐North American Pattern (PNA),” Westward negative “PNA,” and Prominent Alaskan ridge. These four clusters, shown in multiple variables, evolve very differently and have differing links to the Arctic and tropical Pacific regions. Using binary forecast skill measures and a new copula‐based framework for predicting PEx events, we find LSMP indices that are useful predictors of NorCal PEx events, with moisture‐based variables being the best predictors of PEx events at least 6 days before the onset, and the lower atmospheric variables being better than their upper atmospheric counterparts any day in advance tested. To ensure statistical rigor, the LSMPs analyzed here (with the modified acronym) include local tests of both significance and consistency, which are not always featured in the literature on large‐scale meteorological patterns.

54 ENVIRONMENTAL SCIENCES↗

Cosmology with persistent homology: a Fisher forecast

Abstract Persistent homology naturally addresses the multi-scale topological characteristics of the large-scale structure as a distribution of clusters, loops, and voids. We apply this tool to the dark matter halo catalogs from theQuijotesimulations, and build a summary statistic for comparison with the joint power spectrum and bispectrum statistic regarding their information content on cosmological parameters and primordial non-Gaussianity. Through a Fisher analysis, we find that constraints from persistent homology are tighter for 8 out of the 10 parameters by margins of 13–50%. The complementarity of the two statistics breaks parameter degeneracies, allowing for a further gain in constraining power when combined. We run a series of consistency checks to consolidate our results, and conclude that our findings motivate incorporating persistent homology into inference pipelines for cosmological survey data.

Astronomy & Astrophysics↗

A Statistical Evaluation of WRF-LES Trace Gas Dispersion Using Project Prairie Grass Measurements

In recent years, new measurement systems have been deployed to monitor and quantify methane emissions from the natural gas sector. Large-eddy simulation (LES) has complemented measurement campaigns by serving as a controlled environment in which to study plume dynamics and sampling strategies. However, with few comparisons with controlled-release experiments, the accuracy of LES for modeling natural gas emissions is poorly characterized. In this paper, we evaluate LES from the Weather Research and Forecasting (WRF) Model against Project Prairie Grass campaign measurements and surface layer similarity theory. Using WRF-LES, we simulate continuous emissions from 30 near-surface trace gas sources in two stability regimes: strong convection and weak convection. We examine the impact of grid resolutions ranging from 6.25 to 52 m in the horizontal dimension on model results. We evaluate performance in a statistical framework, calculating fractional bias and conducting Welch’s t tests. WRF-LES accurately simulates observed surface concentrations at 100 m and beyond under strong convection; simulated concentrations pass t tests in this region irrespective of grid resolution. However, in weakly convective conditions with strong winds, WRF-LES substantially overpredicts concentrations—the magnitude of fractional bias often exceeds 30%, and all but one t test fails. The good performance of WRF-LES under strong convection correlates with agreement with local free convection theory and a minimal amount of parameterized turbulent kinetic energy. The poor performance under weak convection corresponds to misalignment with Monin–Obukhov similarity theory and a significant amount of parameterized turbulent kinetic energy.

17 WIND ENERGY↗

Constraining primordial non-Gaussianity from the large scale structure two-point and three-point correlation functions

Surveys of cosmological large-scale structure (LSS) are sensitive to the presence of local primordial non-Gaussianity (PNG), and may be used to constrain models of inflation. Local PNG, characterized by f NL ⁠, the amplitude of the quadratic correction to the potential of a Gaussian random field, is traditionally measured from LSS two-point and three-point clustering via the power spectrum and bi-spectrum. We propose a framework to measure f NL using the configuration space two-point correlation function (2pcf) monopole and three-point correlation function (3pcf) monopole of survey tracers. Our model estimates the effect of the scale-dependent bias induced by the presence of PNG on the 2pcf and 3pcf from the clustering of simulated dark matter haloes. We describe how this effect may be scaled to an arbitrary tracer of the cosmological matter density. The 2pcf and 3pcf of this tracer are measured to constrain the value of f NL ⁠. In LSS surveys, the effect of imaging systematics on two-point statistics is often degenerate with the PNG signal. Our proposed model employs three-point statistics primarily to break this degeneracy. Using simulations of luminous red galaxies observed by the Dark Energy Spectroscopic Instrument (DESI), we demonstrate the accuracy and constraining power of our method. Our forecast indicates the ability to constrain f NL to a precision of σf NL ≈ 22 with one year of DESI survey data, as well as the ability to constrain the imaging systematic weights in situ.

early Universe↗

Dark Energy Survey Year 3 results: Simulation-based 𝑤CDM inference from weak lensing and galaxy clustering maps with deep learning: Analysis design

Data-driven approaches using deep learning are emerging as powerful techniques to extract non-Gaussian information from cosmological large-scale structure. Here, this work presents the first simulation-based inference (SBI) pipeline that combines weak lensing and galaxy clustering maps in a realistic Dark Energy Survey Year 3 (DES Y3) configuration and serves as preparation for a forthcoming analysis of the survey data. We develop a scalable forward model based on the CosmoGridV1 suite of N-body simulations to generate over one million self-consistent mock realizations of DES Y3 at the map level. Leveraging this large dataset, we train deep graph convolutional neural networks on the full survey footprint in spherical geometry to learn low-dimensional features that approximately maximize mutual information with target parameters. These learned compressions enable neural density estimation of the implicit likelihood via normalizing flows in a ten-dimensional parameter space spanning cosmological 𝑤CDM, intrinsic alignment, and linear galaxy bias parameters, while marginalizing over baryonic, photometric redshift, and shear bias nuisances. To ensure robustness, we extensively validate our inference pipeline using synthetic observations derived from both systematic contaminations in our forward model and independent Buzzard galaxy catalogs. Our forecasts yield significant improvements in cosmological parameter constraints, achieving 2−3× higher figures of merit in the 𝛺 𝑚 − 𝑆 8 plane relative to our implementation of baseline two-point statistics and effectively breaking parameter degeneracies through probe combination. These results demonstrate the potential of SBI analyses powered by deep learning for upcoming Stage-IV wide-field imaging surveys.

Thomsen, A. [Zurich, ETH] (ORCID:0000000203099021)↗

Anticipating Technical Expertise and Capability Evolution in Research Communities Using Dynamic Graph Transformers

The ability to anticipate global technical expertise and capability evolution trends is essential for national and global security, especially in safety-critical domains such as nuclear nonproliferation (NN) and rapidly emerging fields like artificial intelligence (AI). Here, in this work, we extend traditional statistical relational learning approaches (e.g., link prediction in collaboration networks) and formulate a problem of anticipating technical expertise and capability evolution using dynamic heterogeneous graph representations. We develop novel capabilities to forecast collaboration patterns, authorship behavior, and technical capability evolution at different granularities (e.g., scientist and institution levels) in two distinct research fields. We implement a dynamic graph transformer (DGT) neural architecture, which pushes the state-of-the-art graph neural network models by: 1) forecasting heterogeneous (rather than homogeneous) nodes and edges; and 2) relying on both discrete- and continuous-time inputs. We demonstrate that our DGT models predict collaboration, partnership, and expertise patterns with 0.26, 0.73, and 0.53 mean reciprocal rank values for AI and 0.48, 0.93, and 0.22 for NN domains. DGT model performance exceeds the best-performing static graph baseline models by 30%–80% across AI and NN domains. Our findings demonstrate that DGT models boost inductive task performance when previously unseen nodes appear in the test data for the domains with emerging collaboration patterns (e.g., AI). Specifically, models accurately predict which established scientists will collaborate with early career scientists and vice versa in the AI domain.

97 MATHEMATICS AND COMPUTING↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗

Simulation budgeting for hybrid effective field theories

In this work, we forecast the number of, and requirements on, N-body simulations needed to train hybrid effective field theory (HEFT) emulators for a range of use cases, using a hybrid of HMcode and perturbation theory as a surrogate model. Our accuracy goals, determined with careful consideration of statistical and systematic uncertainties, are 1% accurate in the high-likelihood range of cosmological parameters, and 2% accurate over a broader parameter space volume for k < 1 h Mpc -1 and z < 3. Focusing in part on the 8-parameter w 0 w a CDM+ m ν cosmological model, we find that < 225 simulations are required to meet our error goals over our wide parameter space, including models with rapidly evolving dark energy, given our simulation and emulator recommendations. For a more restricted parameter space volume, as few as 80 simulations are sufficient. We additionally present simulation forecasts for example use cases, and make the code used in our analyses publicly available. These results offer practical guidance for efficient emulator design and simulation budgeting in future cosmological analyses.

cosmological parameters from LSS↗

Modeling the impact of COVID-19 on air quality in southern California: implications for future control policies

Abstract. In response to the coronavirus disease of 2019 (COVID-19), California issued statewide stay-at-home orders, bringing about abrupt and dramatic reductions in air pollutant emissions. This crisis offers us an unprecedented opportunity to evaluate the effectiveness of emission reductions in terms of air quality. Here we use the Weather Research and Forecasting model with Chemistry (WRF-Chem) in combination with surface observations to study the impact of the COVID-19 lockdown measures on air quality in southern California. Based on activity level statistics and satellite observations, we estimate the sectoral emission changes during the lockdown. Due to the reduced emissions, the population-weighted concentrations of fine particulate matter (PM2.5) decrease by 15 % in southern California. The emission reductions contribute 68 % of the PM2.5 concentration decrease before and after the lockdown, while meteorology variations contribute the remaining 32 %. Among all chemical compositions, the PM2.5 concentration decrease due to emission reductions is dominated by nitrate and primary components. For O3 concentrations, the emission reductions cause a decrease in rural areas but an increase in urban areas; the increase can be offset by a 70 % emission reduction in anthropogenic volatile organic compounds (VOCs). These findings suggest that a strengthened control on primary PM2.5 emissions and a well-balanced control on nitrogen oxides and VOC emissions are needed to effectively and sustainably alleviate PM2.5 and O3 pollution in southern California.

54 ENVIRONMENTAL SCIENCES↗

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models↗

Robust Medium-Voltage Distribution System State Estimation using Multi-Source Data

Due to the lack of sufficient online measurements for distribution system observability, pseudo-measurements from short-term load or distributed renewable energy resources (DERs) forecasting are used. However, the accuracy of them is low and thus significantly limits the performance of distribution system state estimation (DSSE). In this paper, a robust DSSE that integrates multi-source measurement data is proposed. Specifically, the historical low-voltage (LV) side smart meters are used to forecast load and DERs injections via the support vector machine (SVM) with optimally tuned parameters. By contrast, the online smart meters at LV side are utilized to derive equivalent power injections at the MV/LV transformers, yielding more accurate pseudo-measurements compared to the forecasted injections. Furthermore, to deal with bad data caused by communication loss, instrumental errors and cyber attacks, robust DSSE that relies on generalized maximum-likelihood (GM)-estimation criterion is developed. The projection statistics are developed to adjust the weights of each measurement, leading to better balance between pseudo- and real-time measurements. Numerical results conducted on modified IEEE 33-bus system with DG integration demonstrate the effectiveness and robustness of the proposed method.

distribution system state estimation↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

Cluster cosmology without cluster finding

ABSTRACT We propose that observations of supermassive galaxies contain cosmological statistical constraining power similar to conventional cluster cosmology, and we provide promising indications that the associated systematic errors are comparably easier to control. We consider a fiducial spectroscopic and stellar mass complete sample of galaxies drawn from the Dark Energy Spectroscopic Instrument (DESI) and forecast how constraints on Ωm–σ8 from this sample will compare with those from number counts of clusters based on richness λ. At fixed number density, we find that massive galaxies offer similar constraints to galaxy clusters. However, a mass-complete galaxy sample from DESI has the potential to probe lower halo masses than standard optical cluster samples (which are typically limited to λ ≳ 20 and Mhalo ≳ 1013.5 M⊙ h−1); additionally, it is straightforward to cleanly measure projected galaxy clustering wp for such a DESI sample, which we show can substantially improve the constraining power on Ωm. We also compare the constraining power of M*-limited samples to those from larger but mass-incomplete samples [e.g. the DESI Bright Galaxy Survey (BGS) sample]; relative to a lower number density M*-limited samples, we find that a BGS-like sample improves statistical constraints by 60 per cent for Ωm and 40 per cent for σ8, but this uses small-scale information that will be harder to model for BGS. Our initial assessment of the systematics associated with supermassive galaxy cosmology yields promising results. The proposed samples have a ∼10 per cent satellite fraction, but we show that cosmological constraints may be robust to the impact of satellites. These findings motivate future work to realize the potential of supermassive galaxies to probe lower halo masses than richness-based clusters and to potentially avoid persistent systematics associated with optical cluster finding.

(cosmology): large-scale structure of Universe↗

A Nowcasting Approach for Low-Earth-Orbiting Hyperspectral Infrared Soundings within the Convective Environment

Low-Earth-orbiting (LEO) hyperspectral infrared (IR) sounders have significant yet untapped potential for characterizing thermodynamic environments of convective initiation and ongoing convection. While LEO soundings are of value to weather forecasters, the temporal resolution needed to resolve the rapidly evolving thermodynamics of the convective environment is limited. Here, we have developed a novel nowcasting methodology to extend snapshots of LEO soundings forward in time up to 6 h to create a product available within National Weather Service systems for user assessment. Our methodology is based on parcel forward-trajectory calculations from the satellite-observing time to generate future soundings of temperature (T) and specific humidity (q) at regularly gridded intervals in space and time. The soundings are based on NOAA-Unique Combined Atmospheric Processing System (NUCAPS) retrievals from the Suomi National Polar-Orbiting Partnership (Suomi NPP) and NOAA-20 satellite platforms. The tendencies of derived convective available potential energy (CAPE) and convective inhibition (CIN) are evaluated against gridded, hourly accumulated rainfall obtained from the Multi-Radar Multi-Sensor (MRMS) observations for 24 hand-selected cases over the contiguous United States. Areas with forecast increases in CAPE (reduced CIN) are shown to be associated with areas of precipitation. The increases in CAPE and decreases in CIN are largest for areas that have the heaviest precipitation and are statistically significant compared to areas without precipitation. These results imply that adiabatic parcel advection of LEO satellite sounding snapshots forward in time are capable of identifying convective initiation over an expanded temporal scale compared to soundings used only during the LEO satellite overpass time.

54 ENVIRONMENTAL SCIENCES↗

helios: An R package to process heating and cooling degrees for GCAM

helios is an open-source R package that estimates population-weighted heating and cooling degree-hours (HDH and CDH) and degree-days (HDD and CDD) at various temporal (e.g., energy dispatch segments, monthly, yearly) and spatial scales (e.g., U.S. states, global political regions, countries). The degree hour and degree day outputs from helios are used to inform electricity demand load in the Global Change Analysis Model (GCAM) as well as in GCAM-USA (which is the version of GCAM with U.S. state-level details). helios uses a workflow with four steps: processing raw data; calculating heating and cooling degrees; visualizing performance diagnostics; and outputing results in various formats. There are two sources of widely-used climate data compatible with helios: (1) hourly climate data with 12-km resolution that are dynamically downscaled with the Weather Research and Forecasting (WRF) model and projected using a thermal global warming (TGW) approach; and (2) daily climate data with 0.5-degree resolution from the Coupled Model Intercomparison Project (CMIP) that is bias-adjusted and statistical downscaled by the Inter-Sectoral Impact Model Intercomparison Project (ISIMIP). In summary, helios is a model that standardizes methodology of heating and cooling degrees-hours and degree-days using publicly available data and advance the understanding of the impact of spatial and temporal temperature variability on building energy services.

97 MATHEMATICS AND COMPUTING↗