Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Paw-Net: Stacking ensemble deep learning for segmenting scanning electron microscopy images of fine-grained shale samples

Segmentation of scanning electron microscopy (SEM) images is critical yet time-consuming for geological analyses, as it needs to differentiate the boundaries for different mineral objects to facilitate subsequent analyses, such as porosity calculation. Recently, various machine learning methods, especially convolutional neural networks (CNNs), have been explored to segment SEM images of fine-grained shale samples. However, we found that general CNNs do not yield optimal performance due to insufficient training data and imbalanced objects in SEM images. This work has revised the U-Net architecture, a popular approach for biomedical image analyses, by incorporating a loss function that addresses the imbalance issue. Furthermore, we used the ensemble learning method to train multiple models and combined the results to improve the overall performance of segmentation. We prepared 2162 sub-images from raw SEM images in our experiments and divided them into training, validation, and testing datasets. The overall results show that our method improves the average Intersection over Union (IOU) of mineral objects from 0.49 to 0.58, compared to the original U-Net model. Our method can clearly distinguish each object from others with boundaries, even in highly imbalanced images. Training our models takes less than three minutes using a single GPU, while manual labeling can take up to three hours for each image. Furthermore, the method helps geoscientists gain insights quickly and effectively by building neural network models from a small dataset of SEM images.

58 GEOSCIENCES↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

Long short-term memory embedded nudging schemes for nonlinear data assimilation of geophysical flows

Reduced rank nonlinear filters are increasingly utilized in data assimilation of geophysical flows, but often require a set of ensemble forward simulations to estimate forecast covariance. On the other hand, predictor-corrector type nudging approaches are still attractive due to their simplicity of implementation when more complex methods need to be avoided. However, optimal estimate of nudging gain matrix might be cumbersome. In this paper, we put forth a fully nonintrusive recurrent neural network approach based on a long short-term memory (LSTM) embedding architecture to estimate the nudging term, which plays a role not only to force the state trajectories to the observations but also acts as a stabilizer. Furthermore, our approach relies on the power of archival data and the trained model can be retrained effectively due to power of transfer learning in any neural network applications. In order to verify the feasibility of the proposed approach, we perform twin experiments using Lorenz 96 system. Our results demonstrate that the proposed LSTM nudging approach yields more accurate estimates than both extended Kalman filter (EKF) and ensemble Kalman filter (EnKF) when only sparse observations are available. With the availability of emerging AI-friendly and modular hardware technologies and heterogeneous computing platforms, we articulate that our simplistic nudging framework turns out to be computationally more efficient than either the EKF or EnKF approaches.

42 ENGINEERING↗

Bias-Adjustment Methods for Future Subdaily Precipitation Extremes Consistent Across Durations

Model output from climate projections often requires bias-adjustment to compensate for systematic model errors. A bias-adjustment method for extreme precipitation intensity is proposed that preserves the scaling equation for different accumulation levels from hourly to daily, using intensity-duration-frequency (IDF) modeling. A validation is performed within a pseudo-reality setting, based on hourly precipitation from 28 regional climate model projections of the EURO-CORDEX ensemble over Belgium. The scaling-based adjustment methods improve upon previous methods, an optimal method is identified, and, analytical quantile mapping methods must be avoided due to three identified problems. The ensemble mean of the adjusted extreme precipitation intensity obeys the above-mentioned scale-invariance property, which is consistent with observed extreme intensities. We thus show that IDF modeling provides added value in the context of bias-adjustment, and, that the particular IDF model proposed balances well between accuracy and the preservation of desired properties such as scale invariance and consistency among rainfall durations.

54 ENVIRONMENTAL SCIENCES↗

Improving the accuracy of freight mode choice models: A case study using the 2017 CFS PUF data set and ensemble learning techniques

Here, the US Census Bureau has collected two rounds of experimental data from the Commodity Flow Survey, providing shipment-level characteristics of nationwide commodity movements, published in 2012 (i.e., Public Use Microdata) and in 2017 (i.e., Public Use File). With this information, data-driven methods have become increasingly valuable for understanding detailed patterns in freight logistics. In this study, we used the 2017 Commodity Flow Survey Public Use File data set to explore building a high-performance freight mode choice model, considering three main improvements: (1) constructing local models for each separate commodity/industry category; (2) extracting useful geographical features, particularly the derived distance of each freight mode between origin/destination zones; and (3) applying additional ensemble learning methods such as stacking or voting to combine results from local and unified models for improved performance. The proposed method achieved over 92% accuracy without incorporating external information, an over 19% increase compared to directly fitting Random Forests models over 10,000 samples. Furthermore, SHAP (Shapely Additive Explanations) values were computed to explain the outputs and major patterns obtained from the proposed model. The model framework could enhance the performance and interpretability of existing freight mode choice models.

42 ENGINEERING↗

Modeling the Electronic Absorption Spectra of the Indocarbocyanine Cy3

Accurate modeling of optical spectra requires careful treatment of the molecular structures and vibronic, environmental, and thermal contributions. The accuracy of the computational methods used to simulate absorption spectra is limited by their ability to account for all the factors that affect the spectral shapes and energetics. The ensemble-based approaches are widely used to model the absorption spectra of molecules in the condensed-phase, and their performance is system dependent. The Franck–Condon approach is suitable for simulating high resolution spectra of rigid systems, and its accuracy is limited mainly by the harmonic approximation. In this work, the absorption spectrum of the widely used cyanine Cy3 is simulated using the ensemble approach via classical and quantum sampling, as well as, the Franck–Condon approach. The factors limiting the ensemble approaches, including the sampling and force field effects, are tested, while the vertical and adiabatic harmonic approximations of the Franck–Condon approach are also systematically examined. Our results show that all the vertical methods, including the ensemble approach, are not suitable to model the absorption spectrum of Cy3, and recommend the adiabatic methods as suitable approaches for the modeling of spectra with strong vibronic contributions. We find that the thermal effects, the low frequency modes, and the simultaneous vibrational excitations have prominent contributions to the Cy3 spectrum. The inclusion of the solvent stabilizes the energetics significantly, while its negligible effect on the spectral shapes aligns well with the experimental observations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Estimating the CO 2 Fertilization Effect on Extratropical Forest Productivity From Flux‐Tower Observations

Abstract The land sink of anthropogenic carbon emissions, a crucial component of mitigating climate change, is primarily attributed to the CO 2 fertilization effect on global gross primary productivity (GPP). However, direct observational evidence of this effect remains scarce, hampered by challenges in disentangling the CO 2 fertilization effect from other long‐term confounding drivers, particularly climatic changes. Here, we introduce a novel statistical approach to separate the CO 2 fertilization effect on photosynthetic carbon uptake using eddy covariance (EC) records across 38 extratropical forest sites. We find the median stimulation rate of GPP to be 3.2 ± 0.9 gC m −2 yr −1 ppm −1 (or 16.4 ± 4.2% per 100 ppm) under increasing atmospheric CO 2 across these sites, respectively. To validate the robustness of our findings, we test our statistical method using factorial simulations of an ensemble of process‐based land surface models. We address additional factors, including nitrogen deposition and land management, that may impact plant productivity, potentially confounding the attribution to the CO 2 fertilization effect. Assuming these site‐specific effects offset to some extent across sites as random factors, the estimated median value still reflects the strength of the CO 2 fertilization effect. However, disentanglement of these long‐term effects, often inseparable by timescale, requires further causal research. Our study provides direct evidence that the photosynthetic stimulation is maintained under long‐term CO 2 fertilization across multiple EC sites. Such observation‐based quantification is key to constraining the long‐standing uncertainties in the land carbon cycle under rising CO 2 concentrations.

Environmental Sciences & Ecology↗

Persistent and partially mobile oxygen vacancies in Li-rich layered oxides

Increasing the energy density of layered oxide battery electrodes is challenging as accessing high states of delithiation often triggers voltage degradation and oxygen release. Here we utilize transmission-based X-ray absorption spectromicroscopy and ptychography on mechanically cross-sectioned Li 1.18–x Ni 0.21 Mn 0.53 Co 0.08 O 2–δ electrodes to quantitatively profile the oxygen deficiency over cycling at the nanoscale. The oxygen deficiency penetrates into the bulk of individual primary particles (~200 nm) and is well-described by oxygen vacancy diffusion. Using an array of characterization techniques, we demonstrate that, surprisingly, bulk oxygen vacancies that persist within the native layered phase are indeed responsible for the observed spectroscopic changes. We additionally show that the arrangement of primary particles within secondary particles (~5 μm) causes considerable heterogeneity in the extent of oxygen release between primary particles. Finally, our work merges an ensemble of length-spanning characterization methods and informs promising approaches to mitigate the deleterious effects of oxygen release in lithium-ion battery electrodes.

25 ENERGY STORAGE↗

Phase identification using co‐association matrix ensemble clustering

Calibrating distribution system models to aid in the accuracy of simulations such as hosting capacity analysis is increasingly important in the pursuit of the goal of integrating more distributed energy resources. The recent availability of smart meter data is enabling the use of machine learning tools to automatically achieve model calibration tasks. This research focuses on applying machine learning to the phase identification task, using a co‐association matrix‐based, ensemble spectral clustering approach. The proposed method leverages voltage time series from smart meters and does not require existing or accurate phase labels. This work demonstrates the success of the proposed method on both synthetic and real data, surpassing the accuracy of other phase identification research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Effective optimization of atomic decoration in giant and superstructurally ordered crystals with machine learning

Crystals with complicated geometry are often observed with mixed chemical occupancy among Wyckoff sites, presenting a unique challenge for accurate atomic modeling. Similar systems possessing exact occupancy on all the sites can exhibit superstructural ordering, dramatically inflating the unit cell size. In this work, a crystal graph convolutional neural network (CGCNN) is used to predict optimal atomic decorations on fixed crystalline geometries. This is achieved with a site permutation search (SPS) optimization algorithm based on Monte Carlo moves combined with simulated annealing and basin-hopping techniques. Our approach relies on the evidence that, for a given chemical composition, a CGCNN estimates the correct energetic ordering of different atomic decorations, as predicted by electronic structure calculations. This provides a suitable energy landscape that can be optimized according to site occupation, allowing the prediction of chemical decoration in crystals exhibiting mixed or disordered occupancy, or superstructural ordering. Verification of the procedure is carried out on several known compounds, including the superstructurally ordered clathrate compound Rb8Ga27Sb16 and vacancy-ordered perovskite Cs2SnI6, neither of which was previously seen during the neural network training. In addition, the critical temperature of an order–disorder phase transition in solid solution CuZn is probed with our SPS routines by sampling site configuration trajectories in the canonical ensemble. This strategy provides an accurate method for determining favorable decoration in complex crystals and analyzing site occupation at unprecedented speed and scale.

Chemistry↗

Assessing the Sensitivity of the Tropical Cyclone Boundary Layer to the Parameterization of Momentum Flux in the Community Earth System Model

Recent studies have demonstrated that high-resolution (~25 km) Earth System Models (ESMs) have the potential to skillfully predict tropical cyclone (TC) occurrence and intensity. However, biases in ESM TCs still exist, largely due to the need to parameterize processes such as boundary layer (PBL) turbulence. Building on past studies, we hypothesize that the depiction of the TC PBL in ESMs is sensitive to the configuration of the PBL parameterization scheme, and that the targeted perturbation of tunable parameters can reduce biases. The Morris one-at-a-time (MOAT) method is implemented to assess the sensitivity of the TC PBL to tunable parameters in the PBL scheme in an idealized configuration of the Community Atmosphere Model, version 6 (CAM6). The MOAT method objectively identifies several parameters in an experimental version of the Cloud Layers Unified by Binormals (CLUBB) scheme that appreciably influence the structure of the TC PBL. We then perturb the parameters identified by the MOAT method within a suite of CAM6 ensemble simulations and find a reduction in model biases compared to observations and a high-resolution, cloud-resolving model. Importantly, we demonstrate that the high-sensitivity parameters are tied to PBL processes that reduce turbulent mixing and effective eddy diffusivity, and that in CAM6 these parameters alter the TC PBL in a manner consistent with past modeling studies. In this way, we provide an initial identification of process-based input parameters that, when altered, have the potential to improve TC predictions by ESMs.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Assisted Safety Modeling and Analysis of Advanced Reactors

With the advances in computational power and numerical methods, analysts can now rely on first-principle simulations to predict ultra-fine details in a variety of applications. Advances in machine learning (ML) have produced algorithms that can now learn high-level abstractions via hierarchical models. This project aims to leverage advances in ML techniques and the available high-resolution simulation data to develop a novel modeling and simulation (M\&S) methodology for reactor safety analysis. While application-agnostic ML techniques are available, complex physics constraints need to be incorporated into ML techniques to build ML-based closures for computationally efficient predictive simulations. This project intends to develop a physics-guided data-driven multi-scale methodology for M\&S of advanced reactors. The project focuses on thermal fluid (T/F) phenomena, which play major roles in advanced reactor safety. Specifically, we propose a data-driven coarse-mesh turbulence model based on local flow features for the transient analysis of thermal mixing and stratification in a sodium-cooled fast reactor (SFR). The model has a coarse-mesh setup to ensure computational efficiency, while it is trained by fine-mesh computational fluid dynamics (CFD) data with Reynolds-averaged Navier-Stokes (RANS) turbulence model to ensure accuracy. Three different neural networks are developed and tested for loss-of-flow transients in the hot pool of SFR, i.e. the densely connected convolutional neural network (DCNN), long-short-term-memory network based on proper orthogonal decomposition (POD-LSTM), and the DCNN informed by LSTM (DCNN-LSTM). The performances of these three neural networks are evaluated based on baseline models. The DCNN-LSTM model has been chosen for further hyperparameter optimization. Furthermore, based on a simplified two-dimensional case, uncertainty quantification (UQ) of the developed ML-based closure are investigated with three methods, i.e. Monte Carlo dropout, deep ensemble, and Bayesian neural network. The developed ML-based turbulent viscosity closure relation based on deep ensemble is then integrated into the system analysis module SAM and serves as a term in the conservation equations. Such a SAM-ML based procedure guarantees that the obtained results are consistent with the physical constraints of the thermal-fluid system. The SAM-ML simulation on the same loss-of-flow transient showed comparable accuracy with the CFD simulation but with a much coarser mesh setup. Last but not least, the ML-based closure improvement with the support of higher-fidelity data from large eddy simulation (LES) is discussed. As a first step towards this direction, a baseline LES simulation is performed to obtain comparable data with RANS results. Based on the early results, future investigation on further improving the ML-based closure is discussed. We believe the developed approach that combines scientific machine learning with nuclear system analysis code can benefit the advanced reactor community as more accurate safety analyses will better characterize reactor safety margins and reduce licensing efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Database for LDV signal processor performance analysis

A technique for the direct comparison of LDV signal processors is developed, based on the use of a data base of digitized signal bursts obtained from an LDV under various configurations. This data base can be used to evaluate the response of signal processors and processor algorithms to specific signal characteristics and not to generalized simplistic waveforms. Examples from such a data base are presented to illustrate the capabilities of the proposed method. This data base includes signal ensembles obtained with three laser power settings at two transmitted focal lengths.

Baker, Glenn D.↗

Advances in Hyperspectral Image Classification Methods for Vegetation and Agricultural Cropland Studies

Hyperspectral data are becoming more widely available via sensors on airborne and unmanned aerial vehicle (UAV) platforms, as well as proximal platforms. While space-based hyperspectral data continue to be limited in availability, multiple spaceborne Earth-observing missions on traditional platforms are scheduled for launch, and companies are experimenting with small satellites for constellations to observe the Earth, as well as for planetary missions. Land cover mapping via classification is one of the most important applications of hyperspectral remote sensing and will increase in significance as time series of imagery are more readily available. However, while the narrow bands of hyperspectral data provide new opportunities for chemistry-based modeling and mapping, challenges remain. Hyperspectral data are high dimensional, and many bands are highly correlated or irrelevant for a given classification problem. For supervised classification methods, the quantity of training data is typically limited relative to the dimension of the input space. The resulting Hughes phenomenon, often referred to as the curse of dimensionality, increases potential for unstable parameter estimates, overfitting, and poor generalization of classifiers. This is particularly problematic for parametric approaches such as Gaussian maximum likelihood–based classifiers that have been the backbone of pixel-based multispectral classification methods. This issue has motivated investigation of alternatives, including regularization of the class covariance matrices, ensembles of weak classifiers, development of feature selection and extraction methods, adoption of nonparametric classifiers, and exploration of methods to exploit unlabeled samples via semi-supervised and active learning. Data sets are also quite large, motivating computationally efficient algorithms and implementations. This chapter provides an overview of the recent advances in classification methods for mapping vegetation using hyperspectral data. Three data sets that are used in the hyperspectral classification literature (e.g., Botswana Hyperion satellite data and AVIRIS airborne data over both Kennedy Space Center and Indian Pines) are described in Section 3.2 and used to illustrate methods described in the chapter. An additional high-resolution hyperspectral data set acquired by a SpecTIR sensor on an airborne platform over the Indian Pines area is included to exemplify the use of new deep learning approaches, and a multiplatform example of airborne hyperspectral data is provided to demonstrate transfer learning in hyperspectral image classification. Classical approaches for supervised and unsupervised feature selection and extraction are reviewed in Section 3.3. In particular, nonlinearities exhibited in hyperspectral imagery have motivated development of nonlinear feature extraction methods in manifold learning, which are outlined in Section 3.3.1.4. Spatial context is also important in classification of both natural vegetation with complex textural patterns and large agricultural fields with significant local variability within fields. Approaches to exploit spatial features at both the pixel level (e.g., co-occurrence–based texture and extended morphological attribute profiles [EMAPs]) and integration of segmentation approaches (e.g., HSeg) are discussed in this context in Section 3.3.2. Recently, classification methods that leverage nonparametric methods originating in the machine learning community have grown in popularity. An overview of both widely used and newly emerging approaches, including support vector machines (SVMs), Gaussian mixture models, and deep learning based on convolutional neural networks is provided in Section 3.4. Strategies to exploit unlabeled samples, including active learning and metric learning, which combine feature extraction and augmentation of the pool of training samples in an active learning framework, are outlined in Section 3.5. Integration of image segmentation with classification to accommodate spatial coherence typically observed in vegetation is also explored, including as an integrated active learning system. Exploitation of multisensor strategies for augmenting the pool of training samples is investigated via a transfer learning framework in Section 3.5.1.2. Finally, we look to the future, considering opportunities soon to be provided by new paradigms, as hyperspectral sensing is becoming common at multiple scales from ground-based and airborne autonomous vehicles to manned aircraft and space-based platforms.

Pasolli, Edoardo↗

Efficient Distance-based Global Sensitivity Analysis for Terrestrial Ecosystem Modeling

Sensitivity analysis in terrestrial ecosystem modeling is important for understanding controlling processes, guiding model development, and targeting new observations to reduce parameter and prediction uncertainty. Complex and computationally expensive terrestrial ecosystem models (TEM) limit the number of ensemble simulations, requiring sophisticated and efficient methods to analyze sensitivities of multiple model responses to different types of parameter uncertainties. In this study, we propose a distance-based global sensitivity analysis (DGSA) method. DGSA first classifies model response samples into a small set of discrete classes and then calculates the distance between parameter frequency distributions in different classes to measure the parameter sensitivity. The principle is that, if the parameter distribution is the same in each class, then the model response is insensitive to the parameter, while a large difference in the distributions indicates the parameter is influential to the response. Built on this idea, DGSA can be applied to analyze sensitivity of a single and a group of responses to different kinds of parameter uncertainties including continuous, discrete and even stochastic. Besides the main-effect sensitivity from a single parameter, DGSA can also quantify the sensitivity from parameter interactions. Additionally, DGSA is computationally efficient which can use a small number of model evaluations to obtain an accurate and statistically significant result. We applied DGSA to two TEMs, one having eight parameters and three kinds of model responses, and the other having 47 parameters and a long-period response. We demonstrated that DGSA can be used for sensitivity problems with multiple responses and high-dimensional parameters efficiently.

Lu, Dan↗

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Unitary Designs from Random Sums and Permutations

A unitary k-design is an ensemble of unitaries that matches the first k moments of the Haar measure. In this work, we provide two efficient constructions of k-designs on n-qubits using new random matrix theory techniques. Our first construction is based on exponentiating sums of random i.i.d. Hermitian matrices and uses O(k2n2)-many gates. In the spirit of central limit theorems, we show that this random sum approximates the Gaussian Unitary Ensemble (GUE). We then show that the product of just two exponentiated GUE matrices is already approximately Haar random. Our second construction is based on products of exponentiated sums of random permutations and uses Õ(k poly (n)) many gates. The k dependence is optimal (up to polylogarithmic factors) and is inherited from the efficiency of existing k-wise independent permutations. Furthermore, replacing random permutations with quantum-secure pseudorandom permutations (PRPs), we also obtain a pseudorandom unitary (PRU) ensemble that is secure under nonadaptive queries. A central feature of both proofs is a new connection between the polynomial method in quantum query complexity and the large-dimension (N) expansion in random matrix theory. In particular, the first construction uses the polynomial method to control high moments of certain random matrix ensembles without requiring delicate Weingarten calculations. In doing so, we define and solve a moment problem on the unit circle, asking whether a finite number of equally weighted points can reproduce a given set of moments. In our second construction, the key step is to exhibit an orthonormal basis for irreducible representations of the partition algebra that has a low-degree large-N expansion. This allows us to show that the distinguishing probability is a low-degree rational polynomial of the dimension N.

algebra↗