Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical change detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Probabilistic Model-Based Diagnostic Framework for Nuclear Engineering Systems

A fault diagnostic framework was investigated in this study for applications in thermal–hydraulic systems of nuclear power plants. The proposed framework consists of quantitative model-based diagnosis, statistical change detection and probabilistic reasoning. The use of physics-based diagnostic models provides high detection sensitivity and allows noise and measurement uncertainty to be incorporated robustly. Performance-related parametric models for each component are constructed based on first principles. Numerical model residuals are generated using the concept of analytical redundancy. Statistical change detection methods are employed to detect non-zero residuals in the presence of uncertainty. The diagnosis task is performed using Bayesian inference to detect and localize possible faults. Application to a single-phase heat exchanger for demonstration showed that the proposed probabilistic framework can provide improved results in comparison with traditional approaches while remaining less sensitive to false alarms in the presence of measurement and modeling uncertainty.

Bayesian network↗

Toward Statistical Real-Time Power Fault Detection

We propose statistical fault detection methodology based on high-frequency data streams that are becoming available in modern power grids. Our approach can be treated as an online (sequential) change point monitoring methodology. However, due to the mostly unexplored and very nonstandard structure of high-frequency power grid streaming data, substantial new statistical development is required to make this methodology practically applicable. The paper includes development of scalar detectors based on multichannel data streams, determination of data-driven alarm thresholds and investigation of the performance and robustness of the new tools. Due to a reasonably large database of faults, we can calculate frequencies of false and correct fault signals, and recommend implementations that optimize these empirical success rates.

bolted faults↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Search for a Cloud Phase Feedback in the Arctic Climate System

This project was motivated by a hypothesis involving the transition in lower troposphere temperatures across the freezing point of water. Specifically, at temperatures just below to about ten degrees below freezing, Arctic clouds should be in a mixed-phase, with strong influences from secondary ice production (e.g., the Hallett-Mossop process). At temperatures just above freezing, clouds should become entirely deglaciated. We hypothesized that, with multiple years of ARM data, a statistically significant change could be detected in cloud radiative properties and surface radiative fluxes that could be directly attributed to this phase change. Furthermore, in a gradually warming climate, these phase change-related responses in surface radiation would represent a cloud phase feedback as part of Arctic amplification. We designed this research project to coincide with a new ARM Arctic cloud radar product developed by Ed Luke and collaborators at BNL (Luke et al., 2021: PNAS, doi:10.1073/pnas.2021387118) that explicitly contains SIP cloud properties retrieved from radar data. This research project has successfully concluded after analysis of a much larger North Slope of Alaska (NSA) data sample than originally anticipated, and the results are quite different than originally hypothesized.

58 GEOSCIENCES↗

Evaluation of precipitation indices in suites of dynamically and statistically downscaled regional climate models over Florida

Abstract The present work evaluates historical precipitation and its indices defined by the Expert Team on Climate Change Detection and Indices (ETCCDI) in suites of dynamically and statistically downscaled regional climate models (RCMs) against NOAA’s Global Historical Climatology Network Daily (GHCN-Daily) dataset over Florida. The models examined here are: (1) nested RCMs involved in the North American CORDEX (NA-CORDEX) program, (2) variable resolution Community Earth System Models (VR-CESM), (3) Coupled Model Intercomparison Project phase 5 (CMIP5) models statistically downscaled using localized constructed analogs (LOCA) technique. To quantify observational uncertainty, three in situ-based (PRISM, Livneh, CPC) and three reanalysis (ERA5, MERRA2, NARR) datasets are also evaluated against the station data. The reanalyses and dynamically downscaled RCMs generally underestimate the magnitude of the monthly precipitation and the frequency of the extreme rainfall in summer. The models forced with CanESM2 miss the phase of the seasonality of extreme precipitation. All models and reanalyses severely underestimate both the mean and interannual variability of mean wet-day precipitation (SDII), consecutive dry days (CDD), and overestimate consecutive wet days (CWD). Metric analysis suggests large uncertainty across NA-CORDEX models. Both the LOCA and VR-CESM models perform better than the majority of models. Overall, RegCM4 and WRF models perform poorer than the median model performance. The performance uncertainty across models is comparable to that in the reanalyses. Specifically, NARR performs poorer than the median model performance in simulating the mean indices and MERRA2 performs worse than the majority of models in capturing the interannual variability of the indices.

54 ENVIRONMENTAL SCIENCES↗

Enhanced climate reproducibility testing with false discovery rate correction

Simulating the Earth's climate is an important and complex problem, thus climate models are similarly complex, comprised of millions of lines of code. In order to appropriately utilize the latest computational and software infrastructure advancements in Earth system models running on modern hybrid computing architectures to improve their performance, precision, accuracy, or all three; it is important to ensure that model simulations are repeatable and robust. This introduces the need for establishing statistical or non-bit-for-bit reproducibility, since bit-for-bit reproducibility may not always be achievable. Here, we propose a short-simulation ensemble-based test for an atmosphere model to evaluate the null hypothesis that modified model results are statistically equivalent to that of the original model. We implement this test in version 2 of the US Department of Energy's Energy Exascale Earth System Model (E3SM). The test evaluates a standard set of output variables across the two simulation ensembles and uses a false discovery rate correction to account for multiple testing. The false positive rates of the test are examined using re-sampling techniques on large simulation ensembles and are found to be lower than the currently implemented bootstrapping-based testing approach in E3SM. We also evaluate the statistical power of the test using perturbed simulation ensemble suites, each with a progressively larger magnitude of change to a tuning parameter. The new test is generally found to exhibit more statistical power than the current approach, being able to detect smaller changes in parameter values with higher confidence.

Kelleher, Michael E. [Oak Ridge National Laborator↗

Detection of DoS Attacks Using ARFIMA Modeling of GOOSE Communication in IEC 61850 Substations

Integration of Information and Communication Technology (ICT) in modern smart grids (SGs) offers many advantages including the use of renewables and an effective way to protect, control and monitor the energy transmission and distribution. To reach an optimal operation of future energy systems, availability, integrity and confidentiality of data should be guaranteed. Research on the cyber-physical security of electrical substations based on IEC 61850 is still at an early stage. In the present work, we first model the network traffic data in electrical substations, then, we present a statistical Anomaly Detection (AD) method to detect Denial of Service (DoS) attacks against the Generic Object Oriented Substation Event (GOOSE) network communication. According to interpretations on the self-similarity and the Long-Range Dependency (LRD) of the data, an Auto-Regressive Fractionally Integrated Moving Average (ARFIMA) model was shown to describe well the GOOSE communication in the substation process network. Based on this ARFIMA-model and in view of cyber-physical security, an effective model-based AD method is developed and analyzed. Two variants of the statistical AD considering statistical hypothesis testing based on the Generalized Likelihood Ratio Test (GLRT) and the cumulative sum (CUSUM) are presented to detect flooding attacks that might affect the availability of the data. Our work presents a novel AD method, with two different variants, tailored to the specific features of the GOOSE traffic in IEC 61850 substations. The statistical AD is capable of detecting anomalies at unknown change times under the realistic assumption of unknown model parameters. The performance of both variants of the AD method is validated and assessed using data collected from a simulation case study. We perform several Monte-Carlo simulations under different noise variances. The detection delay is provided for each detector and it represents the number of discrete time samples after which an anomaly is detected. In fact, our statistical AD method with both variants (CUSUM and GLRT) has around half the false positive rate and a smaller detection delay when compared with two of the closest works found in the literature. Our AD approach based on the GLRT detector has the smallest false positive rate among all considered approaches. Whereas, our AD approach based on the CUSUM test has the lowest false negative rate thus the best detection rate. Depending on the requirements as well as the costs of false alarms or missed anomalies, both variants of our statistical detection method can be used and are further analyzed using composite detection metrics.

IEC 61850 electrical substations↗

Sensitive Detection of Structural Differences using a Statistical Framework for Comparative Crystallography

Chemical and conformational changes underlie the functional cycles of proteins. Comparative crystallography can reveal these changes over time, over ligands, and over chemical and physical perturbations in atomic detail. A key difficulty, however, is that the resulting observations must be placed on the same scale by correcting for experimental factors. We recently introduced a Bayesian framework for correcting (scaling) X-ray diffraction data by combining deep learning with statistical priors informed by crystallographic theory. To scale comparative crystallography data, we here combine this framework with a multivariate statistical theory of comparative crystallography. By doing so, we find strong improvements in the detection of protein dynamics, element-specific anomalous signal, and the binding of drug fragments.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U↗

Sensitive detection of structural dynamics using a statistical framework for comparative crystallography

Chemical and conformational changes are crucial to protein function and its pharmacological control. X-ray crystallography can reveal these changes in atomic detail, but standard analysis methods, which refine separate datasets, often overlook differences that are subtle or arise in only a subset of molecules. Direct comparison of crystallographic datasets is, in principle, more powerful, but systematic errors (“scales”) often mask changes in the crystallographic observables (“structure factors”). Machine learning algorithms that jointly estimate scales and structure factors can address this limitation. Here, we augment this approach with multivariate, structured priors derived from crystallographic theory, implemented in the variational deep learning framework Careless. Doing so strongly improves the detection of protein dynamics, element-specific anomalous signals, and the binding of drug candidates, offering a robust approach to comparative crystallography and, potentially, to detection of protein dynamics by other structure determination methods.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U↗

Robust global detection of forced changes in mean and extreme precipitation despite observational disagreement on the magnitude of change

Detection and attribution (D&A) of forced precipitation change are challenging due to internal variability, limited spatial, and temporal coverage of observational records and model uncertainty. These factors result in a low signal-to-noise ratio of potential regional and even global trends. Here, we use a statistical method – ridge regression – to create physically interpretable fingerprints for the detection of forced changes in mean and extreme precipitation with a high signal-to-noise ratio. The fingerprints are constructed using Coupled Model Intercomparison Project phase 6 (CMIP6) multi-model output masked to match coverage of three gridded precipitation observational datasets – GHCNDEX, HadEX3, and GPCC – and are then applied to these observational datasets to assess the degree of forced change detectable in the real-world climate in the period 1951–2020. We show that the signature of forced change is detected in all three observational datasets for global metrics of mean and extreme precipitation. Forced changes are still detectable from changes in the spatial patterns of precipitation even if the global mean trend is removed from the data. This shows the detection of forced change in mean and extreme precipitation beyond a global mean trend is robust and increases confidence in the detection method's power as well as in climate models' ability to capture the relevant processes that contribute to large-scale patterns of change. We also find, however, that detectability depends on the observational dataset used. Not only coverage differences but also observational uncertainty contribute to dataset disagreement, exemplified by the times of emergence of forced change from internal variability ranging from 1998 to 2004 among datasets. Furthermore, different choices for the period over which the forced trend is computed result in different levels of agreement between observations and model projections. These sensitivities may explain apparent contradictions in recent studies on whether models under- or overestimate the observed forced increase in mean and extreme precipitation. Lastly, the detection fingerprints are found to rely primarily on the signal in the extratropical Northern Hemisphere, which is at least partly due to observational coverage but potentially also due to the presence of a more robust signal in the Northern Hemisphere in general.

54 ENVIRONMENTAL SCIENCES↗

Simulated Changes in Tropical Cyclone Size, Accumulated Cyclone Energy and Power Dissipation Index in a Warmer Climate

Detection, attribution and projection of changes in tropical cyclone intensity statistics are made difficult from the potentially decreasing overall storm frequency combined with increases in the peak winds of the most intense storms as the climate warms. Multi-decadal simulations of stabilized climate scenarios from a high-resolution tropical cyclone permitting atmospheric general circulation model are used to examine simulated global changes from warmer temperatures, if any, in estimates of tropical cyclone size, accumulated cyclonic energy and power dissipation index. Changes in these metrics are found to be complicated functions of storm categorization and global averages of them are unlikely to easily reveal the impact of climate change on future tropical cyclone intensity statistics.

Wehner, Michael (ORCID:0000000159910082)↗

Anomaly Detection, Localization and Classification using Drifting Synchrophasor Data Streams

With ongoing automation and digitization of the electric power system, several Phasor Measurement Units(PMUs) have been deployed for monitoring and control. PMU data can have multiple anomalies, and many of the researchers in the past have concentrated on training machine/deep learning algorithms offline for anomaly detection over PMU data (i.e., not in real time). These machine/deep learning algorithms, when trained offline on a sample rather than a population of the dataset, fail to consider the dynamic behavior of the power grid in real-time, resulting in low accuracy. Considering the dynamic behavior of the power grid (e.g., change in load, generation, distributed energy resources (DERs) switching, network, controls), the definition of data anomalies varies in time and requires online training. A fundamental challenge is to enable online (i.e., real-time) training of machine/deep learning algorithms for anomaly detection over streaming PMU data. While machine/deep learning is often desirable to manage data streams, training a deep learning algorithm over streaming PMU data is nontrivial due to changes in data statistics caused by dynamic streaming data. This paper proposes PMUNET: a novel device-level deep learning-based data-driven approach for anomaly detection, localization, and classification over streaming PMU data, using online learning and multivariate data-drift detection algorithm .Two variants of PMUNET, Dynamic data Change Driven Learning (DCDL) and Continuity Driven Learning (CDL), are proposed and compared. DCDL aims to train the deep learning algorithm whenever the definition of anomaly changes due to the power grid dynamics. On the other hand, CDL continuously trains the deep learning algorithm over the PMU data-stream. The experimental results verify that DCDL outperforms CDL and other efficient anomaly detection methods over multiple events such as faults and load/ generator/capacitor/DERs variations/switching for IEEE 14 and 39 Bus test system as well as real PMU industrial data. The result verifies that DCDL variant of PMUNET improves over existing approach with a gain of 2% - 10% in terms of accuracy, false-positive rate, and false-negative rate.

adversarial deep learning↗

The Relationship between Precipitation and Precipitable Water in CMIP6 Simulations and Implications for Tropical Climatology and Change

It is well documented that over the tropical oceans, column-integrated precipitable water (pw) and precipitation (P) have a nonlinear relationship. In this study moisture budget analysis is used to examine this P–pw relationship in a normalized precipitable water framework. It is shown that the parameters of the nonlinear relationship depend on the vertical structure of moisture convergence. Specifically, the precipitable water values at which precipitation is balanced independently by evaporation versus by moisture convergence define a critical normalized precipitable water, pw nc . This is a measure of convective inhibition that separates tropical precipitation into two regimes: a local evaporation-controlled regime with widespread drizzle and a precipitable water–controlled regime. Most of the 17 CMIP6 historical simulations examined here have higher pw nc compared to ERA5, and more frequently they operate in the drizzle regime. When compared to observations, they overestimate precipitation over the high-evaporation oceanic regions off the equator, thereby producing a ‘‘double ITCZ’’ feature, while underestimating precipitation over the large tropical landmasses and over the climatologically moist oceanic regions near the equator. The responses to warming under the SSP585 scenario are also examined using the normalized precipitable water framework. It is shown that the critical normalized precipitable water value at which evaporation versus moisture convergence balance precipitation decreases as a result of the competing dynamic and thermodynamic responses to warming, resulting in an increase in drizzle and total precipitation. Statistically significant historical trends corresponding to the thermodynamic and dynamic changes are detected in ERA5 and in lowintensity drizzle precipitation in the PERSIANN precipitation dataset.

54 ENVIRONMENTAL SCIENCES↗

Cyber-Physical System Implementation for Manufacturing With Analytics in the Cloud Layer

Effective and efficient modern manufacturing operations require the acceptance and incorporation of the fourth industrial revolution, also known as Industry 4.0. Traditional shop floors are evolving their production into smart factories. To continue this trend, a specific architecture for the cyber-physical system is required, as well as a systematic approach to automate the application of algorithms and transform the acquired data into useful information. This work makes use of an approach that distinguishes three layers that are part of the existing Industry 4.0 paradigm: edge, fog, and cloud. Each of the layers performs computational operations, transforming the data produced in the smart factory into useful information. Trained or untrained methods for data analytics can be incorporated into the architecture. A case study is presented in which a real-time statistical control process algorithm based on control charts was implemented. The algorithm automatically detects changes in the material being processed in a computerized numerical control (CNC) machine. The algorithm implemented in the proposed architecture yielded short response times. The performance was effective since it automatically adapted to the machining of aluminum and then detected when the material was switched to steel. The data were backed up in a database that would allow traceability to the line of g-code that performed the machining.

97 MATHEMATICS AND COMPUTING↗

Parsing Weather Variability and Wildfire Effects on the Post-Fire Changes in Daily Stream Flows: A Quantile-Based Statistical Approach and Its Application

Determining wildland fire impacts on streamflow can be problematic as the hydrology in burned watersheds is influenced by post-fire weather conditions. Here, this study presents a quantile-based analytical framework for assessing fire impacts on low and peak daily flow magnitudes, while accounting for post-fire weather influences. This framework entails (a) the bootstrap method to compute the relative change in the post-fire annual flow and weather statistics, (b) double mass analysis to detect if post-fire baseflow and quick-flow yield ratios are significantly altered, and (c) a quantile regression method to parse fire effects on flow at a specific quantile. We illustrate the applicability of this analytical framework using 44 western US streams with at least 5% of their watershed area burned. Results indicate that large, high-severity burns in upland watersheds can raise the streamflow magnitude at the 0. 05 th and 0. 95 th quantiles for at least the five post-fire years. Quantile regression results show that the median fire-related increase in flow for the five post-fire years can be up to 5,000% (Standard Error; S.E. < 2%) at the 0. 05 th E quantile and 161% (S.E. < 10%) at the 0. 95 th quantile. The fire-related increase in flow was often pronounced at the 0. 05 th quantile for streams in the Pacific Northwest and California regions. The difference in fire effects on flow (at both quantiles) across streams was related to post-fire weather, pyrology, physiography, and land cover. The proposed analytical framework can be useful for detecting and quantifying fire effects on the low and peak stream flows in burned watersheds without overlapping disturbances.

54 ENVIRONMENTAL SCIENCES↗

The impact of detection rate changes and correlations on random-coincidence background measurements

Coincidence detection of multiple particles emitted during an experiment can yield a new depth of understanding of the underlying process under study. However, the probability of detecting particles that are generated from the same physical event within a given coincidence time window is generally much lower than that of detecting particles that appear in the same coincidence time window, but were not created from the same physical event, and are therefore detected randomly in coincidence with each other. Thus, accurate and precise methods of measuring this random-coincidence background are essential for a wide variety of fields of science. A method to determine this background directly using the data themselves without any additional experimental run time or fake signals introduced in the data was recently established (O’Donnell, 2016). This method yields a statistical uncertainty on the random-coincidence background that is orders of magnitude smaller than that of the true coincidence data, though the potential for systematic errors of backgrounds from this method was never explored. In this work, we discuss common varieties of correlated and uncorrelated changes in the detection rates of each particle detected in an experiment. Here we demonstrate here that a correlation between particle detection rates from, for example, an incident particle beam that initiates a physical process of interest, creates systematic errors in the random-coincidence background measurement. We also discuss the impact of a variety of other realistic scenarios for rate changes in experiments. Lastly, a method is introduced to correct for errors in the random-coincidence background from any source, yielding an optimization between statistical precision and eliminating potential lingering systematic errors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

RTN-056: Study of the Photon Transfer Curve in the CCD detectors of the Vera C. Rubin Observatory

The RECA internship program provides Colombian students with an opportunity to enhance their research skills in Astronomy, Astrophysics, and Cosmology. During this three-month program, our main objective was to study the Photon Transfer Curves (PTC) of the Vera C. Rubin Observatory, specifically the gain, and to compare it with the gain obtained through pairs of flats. Overall, the study of PTCs is crucial in understanding the performance of detectors and instruments used in Astronomy. The Vera C. Rubin Observatory is an important facility that will enable researchers to carry out a wide range of studies in this field, making it essential to investigate its gain performance. We used run 13144 to construct the PTCs and 13186 to analyze the crosstalk. We employed the LSST Science Pipelines (also known as the DM stack), a software under development for this observatory, which performs all the necessary reductions for the construction of the PTCs. We also used simulations to replicate the observed effects. Initially, we found a 5% difference between the gain calculated by PTC and pairs of flats for a flow range between 5000 and 10000 ADU. Simulations showed that this difference was due to the handling of statistics and the assumption that the distribution following the Lupton equation is Gaussian. We found an error interval for this flow region based on the vendor, with (1.8 ± 0.7, 4.1 ± 0.9) % for E2V and (0.85 ± 0.7, 2.2 ± 0.9) % for ITL. From the PTC, we also obtained the average Full Well Capacity of LSSTCam as 130000 ± 10000$ electrons. We identified a list of segments where we found differences with the results obtained by SLAC National Acceleration Laboratory in PTC parameters, low saturation level, or other defects. We detected and corrected the effect of statistics in the gain calculation using pairs of flats and proposed a code change, which was implemented in the pipeline software. We do not recommend correcting for crosstalk as it does not significantly affect the parameters and does not change the shape of the PTC. However, the opposite is true for the nonlinearity correction.

79 ASTRONOMY AND ASTROPHYSICS↗

Extreme metrics from large ensembles: investigating the effects of ensemble size on their estimates

Abstract. We consider the problem of estimating the ensemble sizes required to characterize the forced component and the internal variability of a number of extreme metrics. While we exploit existing large ensembles, our perspective is that of a modeling center wanting to estimate a priori such sizes on the basis of an existing small ensemble (we assume the availability of only five members here). We therefore ask if such a small-size ensemble is sufficient to estimate accurately the population variance (i.e., the ensemble internal variability) and then apply a well-established formula that quantifies the expected error in the estimation of the population mean (i.e., the forced component) as a function of the sample size n, here taken to mean the ensemble size. We find that indeed we can anticipate errors in the estimation of the forced component for temperature and precipitation extremes as a function of n by plugging into the formula an estimate of the population variance derived on the basis of five members. For a range of spatial and temporal scales, forcing levels (we use simulations under Representative Concentration Pathway 8.5) and two models considered here as our proof of concept, it appears that an ensemble size of 20 or 25 members can provide estimates of the forced component for the extreme metrics considered that remain within small absolute and percentage errors. Additional members beyond 20 or 25 add only marginal precision to the estimate, and this remains true when statistical inference through extreme value analysis is used. We then ask about the ensemble size required to estimate the ensemble variance (a measure of internal variability) along the length of the simulation and – importantly – about the ensemble size required to detect significant changes in such variance along the simulation with increased external forcings. Using the F test, we find that estimates on the basis of only 5 or 10 ensemble members accurately represent the full ensemble variance even when the analysis is conducted at the grid-point scale. The detection of changes in the variance when comparing different times along the simulation, especially for the precipitation-based metrics, requires larger sizes but not larger than 15 or 20 members. While we recognize that there will always exist applications and metric definitions requiring larger statistical power and therefore ensemble sizes, our results suggest that for a wide range of analysis targets and scales an effective estimate of both forced component and internal variability can be achieved with sizes below 30 members. This invites consideration of the possibility of exploring additional sources of uncertainty, such as physics parameter settings, when designing ensemble simulations.

54 ENVIRONMENTAL SCIENCES↗