Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

A comparative study of two-way and offline coupled WRF v3.4 and CMAQ v5.0.2 over the contiguous US: performance evaluation and impacts of chemistry–meteorology feedbacks on air quality

The two-way coupled Weather Research and Forecasting and Community Multiscale Air Quality (WRF-CMAQ) model has been developed to more realistically represent the atmosphere by accounting for complex chemistry–meteorology feedbacks. In this study, we present a comparative analysis of two-way (with consideration of both aerosol direct and indirect effects) and offline coupled WRF v3.4 and CMAQ v5.0.2 over the contiguous US. Long-term (5 years from 2008 to 2012) simulations using WRF-CMAQ with both offline and two-way coupling modes are carried out with anthropogenic emissions based on multiple years of the U.S. National Emission Inventory and chemical initial and boundary conditions derived from an advanced Earth system model (i.e., a modified version of the Community Earth System Model/Community Atmospheric Model). The comprehensive model evaluations show that both two-way WRF-CMAQ and WRF-only simulations perform well for major meteorological variables such as temperature at 2 m, relative humidity at 2 m, wind speed at 10 m, precipitation (except for against the National Climatic Data Center data), and shortwave and longwave radiation. Both two-way and offline CMAQ also show good performance for ozone (O3) and fine particulate matter (PM 2.5 ). Due to the consideration of aerosol direct and indirect effects, two-way WRF-CMAQ shows improved performance over offline coupled WRF and CMAQ in terms of spatiotemporal distributions and statistics, especially for radiation, cloud forcing, O 3 , sulfate, nitrate, ammonium, elemental carbon, tropospheric O 3 residual, and column nitrogen dioxide (NO 2 ). For example, the mean biases have been reduced by more than 10 W m –2 for shortwave radiation and cloud radiative forcing and by more than 2 ppb for max 8 h O 3 . However, relatively large biases still exist for cloud predictions, some PM 2.5 species, and PM 10 that warrant follow-up studies to better understand those issues. The impacts of chemistry–meteorological feedbacks are found to play important roles in affecting regional air quality in the US by reducing domain-average concentrations of carbon monoxide (CO), O 3 , nitrogen oxide (NO x ), volatile organic compounds (VOCs), and PM 2.5 by 3.1 % (up to 27.8 %), 4.2 % (up to 16.2 %), 6.6 % (up to 50.9 %), 5.8 % (up to 46.6 %), and 8.6 % (up to 49.1 %), respectively, mainly due to reduced radiation, temperature, and wind speed. The overall performance of the two-way coupled WRF-CMAQ model achieved in this work is generally good or satisfactory and the improved performance for two-way coupled WRF-CMAQ should be considered along with other factors in developing future model applications to inform policy making.

58 GEOSCIENCES↗

OperonSEQer: A set of machine-learning algorithms with threshold voting for detection of operon pairs using short-read RNA-sequencing data

Operon prediction in prokaryotes is critical not only for understanding the regulation of endogenous gene expression, but also for exogenous targeting of genes using newly developed tools such as CRISPR-based gene modulation. A number of methods have used transcriptomics data to predict operons, based on the premise that contiguous genes in an operon will be expressed at similar levels. While promising results have been observed using these methods, most of them do not address uncertainty caused by technical variability between experiments, which is especially relevant when the amount of data available is small. In addition, many existing methods do not provide the flexibility to determine the stringency with which genes should be evaluated for being in an operon pair. We present OperonSEQer, a set of machine learning algorithms that uses the statistic and p-value from a non-parametric analysis of variance test (Kruskal-Wallis) to determine the likelihood that two adjacent genes are expressed from the same RNA molecule. We implement a voting system to allow users to choose the stringency of operon calls depending on whether your priority is high recall or high specificity. In addition, we provide the code so that users can retrain the algorithm and re-establish hyperparameters based on any data they choose, allowing for this method to be expanded as additional data is generated. We show that our approach detects operon pairs that are missed by current methods by comparing our predictions to publicly available long-read sequencing data. OperonSEQer therefore improves on existing methods in terms of accuracy, flexibility, and adaptability.

59 BASIC BIOLOGICAL SCIENCES↗

A First Principles Approach to Spectral Phonon Transport in Heterostructures

Understanding thermal transport across interfaces which give rise to a thermal resistance (also known as Kapitza resistance) is a critical issue affecting the development of nanotechnologies. Much modern and emergent nanotechnology consist of adjacent materials, and phonon mediated heat transfer governs thermal behavior across internal interfaces in these devices. The physics of thermal transport in solids are governed both by phenomena occurring at the atomic scale and interactions with the material's microstructure. The forecasting of fundamental quantities such as temperature, heat flux and thermal conductivity typically employs the semi-classical Boltzmann transport equation to predict the macroscopic behavior of materials in terms of the microscopic dynamics of its heat carriers. Kapitza resistance was first discovered in liquid helium experiments and has led to a fundamental research thrust in micro and nano-scale heat transport, the behavior of thermal carriers across internal interfaces. Thermal interfacial resistance (TIR) is a widely studied phenomenon, first engaged by Swartz and Pohl through their development of the acoustic and diffuse mismatch methods, then continued through myriad efforts with varying methods and approaches in an attempt to resolve carrier behavior at thermal interfaces. Many of the fundamental approaches to TIR have been at the nanoscale, and research is conducted with molecular dynamics (MD) and density functional theory (DFT) methods. The limitations of these methods is system size; atomistic methods tend to be limited to system sizes of 100,000 atoms or less. Larger length-scale methods have also been pursued, based on the principles of acoustic or diffuse mismatch, but not all include simulation of TIR using a full phonon band spectrum, or temperature dependent methods. Our approach to enabling phonon transport in layered materials draws upon our previous work of demonstrating spectrally coupled phonon transport in homogeneous and heterogeneous materials. We use a semi-analytical approach in which the Bose-Einstein (B-E) statistics set the strength of the phonon radiance in a frequency group, but the B-E statistics are informed with information from the transport system. The B-E statistics in a single frequency group feels the influence of all the groups through the spatial temperature. We also include a new field term which is an indicator of the amount of non-equilibrium behavior of the phonon spectrum---this is added to the phonon source term in all groups to ensure closure and conservation of energy, as the phonon groups in the transport system and the analytical systems are coupled. This work builds upon our previous approach by adding a phonon coupling term at an internal interface, using the principles of the DMM through transmission and reflection coefficients. In this work, the coefficients are determined through computing a common temperature at the interface, influenced by the phonon band structure of both materials, in effect, providing mixing between the two material systems and using the common temperature to set the strength of the phonon radiance at the boundaries on either side of the interface. Our approach uses material properties computed along various crystallographic orientations, and while some isotropy is built into the interface condition, the material properties weight the phonon distributions in the proper crystalline direction. Greater resolution of phonon behavior in proximity to an interface, and more accurate predictions of TIR are obtained. While it is true the assumption of diffuse mismatch can yield inconsistent results compared to experiment especially at low temperatures, this work focuses on room temperature and beyond effects, for future applications in nuclear fuel, or thermoelectric devices; a modified mismatch approach may be feasible if applied properly. Additionally, our methods focus on bridging mesoscale to engineering scale

36 MATERIALS SCIENCE↗

A low-complexity non-intrusive approach to predict the energy demand of buildings over short-term horizons

Reliable, non-intrusive, short-term (of up to 12 hours ahead) prediction of a building's energy demand is a critical component of intelligent energy management applications. A number of such approaches have been proposed over time, utilizing various statistical and, more recently, machine learning techniques, such as decision trees, neural networks and support vector machines. Importantly, all of these works barely outperform simple seasonal auto-regressive integrated moving average models, while their complexity is significantly higher. Here, we propose a novel low-complexity non-intrusive approach that improves the predictive accuracy of the state-of-the-art by up to ~10%. The backbone of our approach is a K-nearest neighbours search method, that exploits the demand pattern of the most similar historical days, and incorporates appropriate time-series pre-processing and easing. In the context of this work, we evaluate our approach against state-of-the-art methods and provide insights on their performance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .

08 HYDROGEN↗

Identifying priority ecosystem services in tidal wetland restoration

Classification systems can be an important tool for identifying and quantifying the importance of relationships, assessing spatial patterns in a standardized way, and forecasting alternative decision scenarios to characterize the potential benefits (e.g., ecosystem services) from ecosystem restoration that improve human health and well-being. We present a top-down approach that systematically leverages ecosystem services classification systems to identify potential services relevant for ecosystem restoration decisions. We demonstrate this approach using the U.S. Environmental Protection Agency’s National Ecosystem Service Classification System Plus (NESCS Plus) to identify those ecosystem services that are relevant to restoration of tidal wetlands. We selected tidal wetland management documents from federal agencies, state agencies, wetland conservation organizations, and land stewards across three regions of the continental United States (northern Gulf of Mexico, Mid-Atlantic, and Pacific Northwest) to examine regional and organizational differences in identified potential benefits of tidal wetland restoration activities and the potential user groups who may benefit. We used an automated document analysis to quantify the frequencies at which different wetland types were mentioned in the management documents along with their associated beneficiary groups and the ecological end products (EEPs) those beneficiaries care about, as defined by NESCS Plus. Results showed that a top combination across all three regions, all four organizations, and all four tidal wetland types was the EEP naturalness paired with the beneficiary people who care (existence). Overall, the Mid-Atlantic region and the land steward organizations mentioned ecosystem services more than the others, and EEPs were mentioned in combination with tidal wetlands as a high-level, more general category than the other more specific tidal wetland types. Certain regional and organizations differences were statistically significant. Those results may be useful in identifying ecosystem services-related goals for tidal wetland restoration. This approach for identifying and comparing ecosystem service priorities is broadly transferrable to other ecosystems or decision-making contexts.

54 ENVIRONMENTAL SCIENCES↗

Observations of the Marine Atmospheric Boundary Layer’s Response to a Solar Eclipse

The atmospheric response to the solar eclipse of 8 April 2024 in North America is investigated with a specific focus on the marine atmospheric boundary layer (MABL). We leverage measurements collected during the Third Wind Forecast Improvement Project (WFIP3), including Doppler lidars, sonic anemometers, and thermodynamic profiler data to investigate the atmospheric response across sites that experienced partial eclipse conditions with nearly 90% obscuration. Using these measurements, we examine eclipse-induced changes in key meteorological parameters, such as temperature, wind speed, and turbulent fluxes. Most previous eclipse studies have been conducted over land, whereas this study provides new observations for both coastal and marine environments, offering additional insight into eclipse-driven variability in the MABL. The findings confirm a notable decrease in downwelling shortwave radiation during the eclipse, which results in rapid cooling of surface air. The temperature reduction ranges from $1.2^\circ \text {C}$ to $1.4^\circ \text {C}$ in coastal regions and from $0.3^\circ \text {C}$ to $0.5^\circ \text {C}$ over the ocean. This analysis suggests that the MABL’s higher thermal inertia compared to coastal regions moderates the temperature decrease during the eclipse. Wind speed exhibits a more complex behavior, as it is influenced by both the MABL and preexisting synoptic conditions. Although a reduction in wind speed is observable up to approximately 140 m above ground level (AGL) at more inland sites, at other locations closer to the coast, this reduction is constrained to the lowest 100 m AGL. Turbulence parameters retrieved from sonic anemometers, such as turbulence kinetic energy, turbulent heat flux, and friction velocity, decrease during the eclipse at coastal sites, accompanied by a brief transition of atmospheric stability from unstable to neutral or weakly stable conditions. For the open-ocean sites, the variability in turbulence statistics and atmospheric stability is minimal during the occurrence of the eclipse.

16 TIDAL AND WAVE POWER↗

Pushing the limits of detectability: mixed dark matter from strong gravitational lenses

One of the frontiers for advancing what is known about dark matter lies in using strong gravitational lenses to characterize the population of the smallest dark matter haloes. There is a large volume of information in strong gravitational lens images – the question we seek to answer is to what extent we can refine this information. To this end, we forecast the detectability of a mixed warm and cold dark matter scenario using the anomalous flux ratio method from strong gravitational lensed images. The halo mass function of the mixed dark matter scenario is suppressed relative to cold dark matter but still predicts numerous low-mass dark matter haloes relative to warm dark matter. Since the strong lensing signal receives a contribution from a range of dark matter halo masses and since the signal is sensitive to the specific configuration of dark matter haloes, not just the halo mass function, degeneracies between different forms of suppression in the halo mass function, relative to cold dark matter, can arise. Here, we find that, with a set of lenses with different configurations of the main deflector and hence different sensitivities to different mass ranges of the halo mass function, the different forms of suppression of the halo mass function between the warm dark matter model and the mixed dark matter model can be distinguished with 40 lenses with Bayesian odds of 30:1.

79 ASTRONOMY AND ASTROPHYSICS↗

Unstructured clinical notes within the 24 hours since admission predict short, mid & long-term mortality in adult ICU patients

Mortality prediction for intensive care unit (ICU) patients is crucial for improving outcomes and efficient utilization of resources. Accessibility of electronic health records (EHR) has enabled data-driven predictive modeling using machine learning. However, very few studies rely solely on unstructured clinical notes from the EHR for mortality prediction. In this work, we propose a framework to predict short, mid, and long-term mortality in adult ICU patients using unstructured clinical notes from the MIMIC III database, natural language processing (NLP), and machine learning (ML) models. Depending on the statistical description of the patients’ length of stay, we define the short-term as 48-hour and 4-day period, the mid-term as 7-day and 10-day period, and the long-term as 15-day and 30-day period after admission. We found that by only using clinical notes within the 24 hours of admission, our framework can achieve a high area under the receiver operating characteristics (AU-ROC) score for short, mid and long-term mortality prediction tasks. The test AU-ROC scores are 0.87, 0.83, 0.83, 0.82, 0.82, and 0.82 for 48-hour, 4-day, 7-day, 10-day, 15-day, and 30-day period mortality prediction, respectively. We also provide a comparative study among three types of feature extraction techniques from NLP: frequency-based technique, fixed embedding-based technique, and dynamic embedding-based technique. Lastly, we provide an interpretation of the NLP-based predictive models using feature-importance scores.

60 APPLIED LIFE SCIENCES↗

Complexity-calibrated benchmarks for machine learning reveal when prediction algorithms succeed and mislead

Abstract Recurrent neural networks are used to forecast time series in finance, climate, language, and from many other domains. Reservoir computers are a particularly easily trainable form of recurrent neural network. Recently, a “next-generation” reservoir computer was introduced in which the memory trace involves only a finite number of previous symbols. We explore the inherent limitations of finite-past memory traces in this intriguing proposal. A lower bound from Fano’s inequality shows that, on highly non-Markovian processes generated by large probabilistic state machines, next-generation reservoir computers with reasonably long memory traces have an error probability that is at least $$\sim 60\%$$ ∼ 60 % higher than the minimal attainable error probability in predicting the next observation. More generally, it appears that popular recurrent neural networks fall far short of optimally predicting such complex processes. These results highlight the need for a new generation of optimized recurrent neural network architectures. Alongside this finding, we present concentration-of-measure results for randomly-generated but complex processes. One conclusion is that large probabilistic state machines—specifically, large $$\epsilon$$ ϵ -machines—are key to generating challenging and structurally-unbiased stimuli for ground-truthing recurrent neural network architectures.

97 MATHEMATICS AND COMPUTING↗

Upstreamness and downstreamness in input–output analysis from local and aggregate information

Abstract Ranking sectors and countries within global value chains is of paramount importance to estimate risks and forecast growth in large economies. However, this task is often non-trivial due to the lack of complete and accurate information on the flows of money and goods between sectors and countries, which are encoded in input–output (I–O) tables. In this work, we show that an accurate estimation of the role played by sectors and countries in supply chain networks can be achieved without full knowledge of the I–O tables, but only relying on local and aggregate information, e.g., the total intermediate demand per sector. Our method, based on a rank-1 approximation to the I–O table, shows consistently good performance in reconstructing rankings (i.e., upstreamness and downstreamness measures for countries and sectors) when tested on empirical data from the world input–output database. Moreover, we connect the accuracy of our approximate framework with the spectral properties of the I–O tables, which ordinarily exhibit relatively large spectral gaps. Our approach provides a fast and analytically tractable framework to rank constituents of a complex economy without the need of matrix inversions and the knowledge of finer intersectorial details.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Using spatio-temporal graph neural networks to estimate fleet-wide photovoltaic performance degradation patterns

Accurate estimation of photovoltaic (PV) system performance is crucial for determining its feasibility as a power generation technology and financial asset. PV-based energy solutions offer a viable alternative to traditional energy resources due to their superior Levelized Cost of Energy (LCOE). A significant challenge in assessing the LCOE of PV systems lies in understanding the Performance Loss Rate (PLR) for large fleets of PV systems. Estimating the PLR of PV systems becomes increasingly important in the rapidly growing PV industry. Precise PLR estimation benefits PV users by providing real-time monitoring of PV module performance, while explainable PLR estimation assists PV manufacturers in studying and enhancing the performance of their products. However, traditional PLR estimation methods based on statistical models have notable drawbacks. Firstly, they require user knowledge and decision-making. Secondly, they fail to leverage spatial coherence for fleet-level analysis. Additionally, these methods inherently assume the linearity of degradation, which is not representative of real world degradation. To overcome these challenges, we propose a novel graph deep learning-based decomposition method called the Spatio-Temporal Graph Neural Network for fleet-level PLR estimation (PV-stGNN-PLR). PV-stGNN-PLR decomposes the power timeseries data into aging and fluctuation components, utilizing the aging component to estimate PLR. PV-stGNN-PLR exploits spatial and temporal coherence to derive PLR estimation for all systems in a fleet and imposes flatness and smoothness regularization in loss function to ensure the successful disentanglement between aging and fluctuation. We have evaluated PV-stGNN-PLR on three simulated PV datasets consisting of 100 inverters from 5 sites. Experimental results show that PV-stGNN-PLR obtains a reduction of 33.9% and 35.1% on average in Mean Absolute Percent Error (MAPE) and Euclidean Distance (ED) in PLR degradation pattern estimation compared to the state-of-the-art PLR estimation methods.

14 SOLAR ENERGY↗

Forecasting the potential of weak lensing magnification to enhance LSST large-scale structure analyses

ABSTRACT Recent works have shown that weak lensing magnification must be included in upcoming large-scale structure analyses, such as for the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), to avoid biasing the cosmological results. In this work, we investigate whether including magnification has a positive impact on the precision of the cosmological constraints, as well as being necessary to avoid bias. We forecast this using an LSST mock catalogue and a halo model to calculate the galaxy power spectra. We find that including magnification has little effect on the precision of the cosmological parameter constraints for an LSST galaxy clustering analysis, where the halo model parameters are additionally constrained by the galaxy luminosity function. In particular, we find that for the LSST gold sample (i < 25.3) including weak lensing magnification only improves the galaxy clustering constraint on Ωm by a factor of 1.03, and when using a very deep LSST mock sample (i < 26.5) by a factor of 1.3. Since magnification predominantly contributes to the clustering measurement and provides similar information to that of cosmic shear, this improvement would be reduced for a combined galaxy clustering and shear analysis. We also confirm that not modelling weak lensing magnification will catastrophically bias the cosmological results from LSST. Magnification must therefore be included in LSST large-scale structure analyses even though it does not significantly enhance the precision of the cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

First detection of the BAO signal from early DESI data

We present the first detection of the baryon acoustic oscillations (BAOs) signal obtained using unblinded data collected during the initial 2 months of operations of the Stage-IV ground-based Dark Energy Spectroscopic Instrument (DESI). From a selected sample of 261 291 luminous red galaxies spanning the redshift interval 0.4 < z < 1.1 and covering 1651 square degrees with a 57.9 per cent completeness level, we report a ∼5σ level BAO detection and the measurement of the BAO location at a precision of 1.7 per cent. Using a bright galaxy sample of 109 523 galaxies in the redshift range 0.1 < z < 0.5, over 3677 square degrees with a 50.0 per cent completeness, we also detect the BAO feature at ∼3σ significance with a 2.6 per cent precision. These first BAO measurements represent an important milestone, acting as a quality control on the optimal performance of the complex robotically actuated, fibre-fed DESI spectrograph, as well as an early validation of the DESI spectroscopic pipeline and data management system. Based on these first promising results, we forecast that DESI is on target to achieve a high-significance BAO detection at sub-per cent precision with the completed 5-yr survey data, meeting the top-level science requirements on BAO measurements. This exquisite level of precision will set new standards in cosmology and confirm DESI as the most competitive BAO experiment for the remainder of this decade.

79 ASTRONOMY AND ASTROPHYSICS↗

Predictive performance of multi-model ensemble forecasts of COVID-19 across European nations

Background: Short-term forecasts of infectious disease burden can contribute to situational awareness and aid capacity planning. Based on best practice in other fields and recent insights in infectious disease epidemiology, one can maximise the predictive performance of such forecasts if multiple models are combined into an ensemble. Here, we report on the performance of ensembles in predicting COVID-19 cases and deaths across Europe between 08 March 2021 and 07 March 2022. Methods: We used open-source tools to develop a public European COVID-19 Forecast Hub. We invited groups globally to contribute weekly forecasts for COVID-19 cases and deaths reported by a standardised source for 32 countries over the next 1–4 weeks. Teams submitted forecasts from March 2021 using standardised quantiles of the predictive distribution. Each week we created an ensemble forecast, where each predictive quantile was calculated as the equally-weighted average (initially the mean and then from 26th July the median) of all individual models’ predictive quantiles. We measured the performance of each model using the relative Weighted Interval Score (WIS), comparing models’ forecast accuracy relative to all other models. We retrospectively explored alternative methods for ensemble forecasts, including weighted averages based on models’ past predictive performance. Results: Over 52 weeks, we collected forecasts from 48 unique models. We evaluated 29 models’ forecast scores in comparison to the ensemble model. We found a weekly ensemble had a consistently strong performance across countries over time. Across all horizons and locations, the ensemble performed better on relative WIS than 83% of participating models’ forecasts of incident cases (with a total N=886 predictions from 23 unique models), and 91% of participating models’ forecasts of deaths (N=763 predictions from 20 models). Across a 1–4 week time horizon, ensemble performance declined with longer forecast periods when forecasting cases, but remained stable over 4 weeks for incident death forecasts. In every forecast across 32 countries, the ensemble outperformed most contributing models when forecasting either cases or deaths, frequently outperforming all of its individual component models. Among several choices of ensemble methods we found that the most influential and best choice was to use a median average of models instead of using the mean, regardless of methods of weighting component forecast models. Conclusions: Our results support the use of combining forecasts from individual models into an ensemble in order to improve predictive performance across epidemiological targets and populations during infectious disease epidemics. Our findings further suggest that median ensemble methods yield better predictive performance more than ones based on means. Our findings also highlight that forecast consumers should place more weight on incident death forecasts than incident case forecasts at forecast horizons greater than 2 weeks. Funding: AA, BH, BL, LWa, MMa, PP, SV funded by National Institutes of Health (NIH) Grant 1R01GM109718, NSF BIG DATA Grant IIS-1633028, NSF Grant No.: OAC-1916805, NSF Expeditions in Computing Grant CCF-1918656, CCF-1917819, NSF RAPID CNS-2028004, NSF RAPID OAC-2027541, US Centers for Disease Control and Prevention 75D30119C05935, a grant from Google, University of Virginia Strategic Investment Fund award number SIF160, Defense Threat Reduction Agency (DTRA) under Contract No. HDTRA1-19-D-0007, and respectively Virginia Dept of Health Grant VDH-21-501-0141, VDH-21-501-0143, VDH-21-501-0147, VDH-21-501-0145, VDH-21-501-0146, VDH-21-501-0142, VDH-21-501-0148. AF, AMa, GL funded by SMIGE - Modelli statistici inferenziali per governare l'epidemia, FISR 2020-Covid-19 I Fase, FISR2020IP-00156, Codice Progetto: PRJ-0695. AM, BK, FD, FR, JK, JN, JZ, KN, MG, MR, MS, RB funded by Ministry of Science and Higher Education of Poland with grant 28/WFSN/2021 to the University of Warsaw. BRe, CPe, JLAz funded by Ministerio de Sanidad/ISCIII. BT, PG funded by PERISCOPE European H2020 project, contract number 101016233. CP, DL, EA, MC, SA funded by European Commission - Directorate-General for Communications Networks, Content and Technology through the contract LC-01485746, and Ministerio de Ciencia, Innovacion y Universidades and FEDER, with the project PGC2018-095456-B-I00. DE., MGu funded by Spanish Ministry of Health / REACT-UE (FEDER). DO, GF, IMi, LC funded by Laboratory Directed Research and Development program of Los Alamos National Laboratory (LANL) under project number 20200700ER. DS, ELR, GG, NGR, NW, YW funded by National Institutes of General Medical Sciences (R35GM119582; the content is solely the responsibility of the authors and does not necessarily represent the official views of NIGMS or the National Institutes of Health). FB, FP funded by InPresa, Lombardy Region, Italy. HG, KS funded by European Centre for Disease Prevention and Control. IV funded by Agencia de Qualitat i Avaluacio Sanitaries de Catalunya (AQuAS) through contract 2021-021OE. JDe, SMo, VP funded by Netzwerk Universitatsmedizin (NUM) project egePan (01KX2021). JPB, SH, TH funded by Federal Ministry of Education and Research (BMBF; grant 05M18SIA). KH, MSc, YKh funded by Project SaxoCOV, funded by the German Free State of Saxony. Presentation of data, model results and simulations also funded by the NFDI4Health Task Force COVID-19 ( https://www.nfdi4health.de/task-force-covid-19-2 ) within the framework of a DFG-project (LO-342/17-1). LP, VE funded by Mathematical and Statistical modelling project (MUNI/A/1615/2020), Online platform for real-time monitoring, analysis and management of epidemic situations (MUNI/11/02202001/2020); VE also supported by RECETOX research infrastructure (Ministry of Education, Youth and Sports of the Czech Republic: LM2018121), the CETOCOEN EXCELLENCE (CZ.02.1.01/0.0/0.0/17-043/0009632), RECETOX RI project (CZ.02.1.01/0.0/0.0/16-013/0001761). NIB funded by Health Protection Research Unit (grant code NIHR200908). SAb, SF funded by Wellcome Trust (210758/Z/18/Z).

60 APPLIED LIFE SCIENCES↗

Assessment of software methods for estimating protein-protein relative binding affinities

A growing number of computational tools have been developed to accurately and rapidly predict the impact of amino acid mutations on protein-protein relative binding affinities. Such tools have many applications, for example, designing new drugs and studying evolutionary mechanisms. In the search for accuracy, many of these methods employ expensive yet rigorous molecular dynamics simulations. By contrast, non-rigorous methods use less exhaustive statistical mechanics, allowing for more efficient calculations. However, it is unclear if such methods retain enough accuracy to replace rigorous methods in binding affinity calculations. This trade-off between accuracy and computational expense makes it difficult to determine the best method for a particular system or study. Here, eight non-rigorous computational methods were assessed using eight antibody-antigen and eight non-antibody-antigen complexes for their ability to accurately predict relative binding affinities (ΔΔG) for 654 single mutations. In addition to assessing accuracy, we analyzed the CPU cost and performance for each method using a variety of physico-chemical structural features. This allowed us to posit scenarios in which each method may be best utilized. Most methods performed worse when applied to antibody-antigen complexes compared to non-antibody-antigen complexes. Rosetta-based JayZ and EasyE methods classified mutations as destabilizing (ΔΔG < -0.5 kcal/mol) with high (83–98%) accuracy and a relatively low computational cost for non-antibody-antigen complexes. Some of the most accurate results for antibody-antigen systems came from combining molecular dynamics with FoldX with a correlation coefficient (r) of 0.46, but this was also the most computationally expensive method. Overall, our results suggest these methods can be used to quickly and accurately predict stabilizing versus destabilizing mutations but are less accurate at predicting actual binding affinities. This study highlights the need for continued development of reliable, accessible, and reproducible methods for predicting binding affinities in antibody-antigen proteins and provides a recipe for using current methods.

59 BASIC BIOLOGICAL SCIENCES↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Are light curve classification metrics good proxies for SN Ia cosmological constraining power?

Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗