Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “online algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

305 records · Page 17

Adaptive model tuning studies for non-invasive diagnostics and feedback control of plasma wakefield acceleration at FACET-II

Beam-driven plasma wakefield acceleration (PWFA) achieves the same energy gain in a single meter, for which conventional accelerators require several kilometers, however much work is still required to match the beam quality of conventional accelerators. The PWFA processes to be studied at the FACET-II facility will utilize extremely short (down to a few fs) high energy (10 GeV), high charge (few nC) electron and eventually positron bunches. The PWFA process is extremely sensitive to the detailed longitudinal current profiles of these bunches and it would be of great benefit to have precise measurement and control of these profiles. We present an adaptive model tuning technique for FACET-II to adaptively tune online models based on real time accelerator and beam data to continuously provide a non-invasive diagnostic which predicts the longitudinal phase space (LPS) of extremely short and highly compressed electron beams, which otherwise require destructive X-band transverse deflecting cavity (XTCAV)-based measurements. Based on simulation studies, our method has the potential for: (1). The development of a non-invasive longitudinal phase space diagnostic by adaptively tuning models based on non-invasive measurements such as energy spread spectra which could be recorded at all of the bunch compressors in the facility. (2). Utilize these diagnostic to perform model-independent feedback-based to achieve desired longitudinal phase space distributions. (3). Utilize the model-independent feedback approach to maximize energy gain while minimizing emittance growth and energy gain variance of the PWFA process by tuning accelerator parameters, while monitoring what longitudinal phase space the algorithm has found, information that will be useful for further analytical and simulation studies.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES↗

Robust Real-Time Modeling of Distribution Systems with Data-Driven Grid-Wise Observability (Final Technical Report)

The overall objective of this project is to leverage existing and emerging sensor measurements to develop data-driven observability enhancement algorithms as well as robust state estimation and parameter identification techniques to enable real-time grid-wise monitoring and modeling of loads and distributed energy resources (DERs). The project resulted in a holistic framework with the following key components: 1) data-driven and machine learning-based grid-edge visibility enhancement, 2) robust branch-current-based state estimation (BCSE), and 3) robust real-time steady-state and dynamic-state modeling of loads and DERs. The research outcomes have provided utility companies better network visibility, higher-fidelity load/DER models, and more accurate assessment of DERs’ impacts, thus, facilitating the integration of renewable energy sources. The synergistic collaborative project among Iowa State University (ISU), Argonne National Laboratory (ANL), Electric Power Research Center (EPRC), SIEMENS Industry, Alliant Energy, Cedar Falls Utilities (CFU), and Maquoketa Valley Electric Cooperative (MVEC) to leverage the team’s extensive expertise and experience in power distribution systems, state estimation, online modeling and identification, and DER integrations. The proposed frameworks have been verified with industry adopted software and attempted integrated with existing tool wherever possible

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Online Tool for Preliminary Design and Techno-Economic Analysis of District Geothermal Heating and Cooling Systems

District geothermal heating and cooling systems (DGHCS) have significant benefits for reducing energy consumption as well as building- and grid-level peak electric demand. Currently, no publicly available tools are available to effectively design and conduct techno-economic analysis of DGHCS. GeoWISE was originally developed for preliminary design and techno-economic analysis of geothermal heating and cooling systems in an individual commercial or residential building. This paper introduces recent upgrades of GeoWISE that allow users to design and conduct techno-economic analysis of DGHCS. Several new features are implemented in GeoWISE to allow selection and specification of multiple new or existing buildings. A database of information for over 125 million existing U.S. buildings was used in GeoWISE that allows users easily locate existing buildings of interest based on street addresses, and optionally edit information of the buildings (e.g., footprint, vintage, principal functions, number of floors, window-to-wall ratio). Unique energy simulation models of the selected buildings are then automatically created using the Automatic Building Energy Modeling (AutoBEM) and EnergyPlus simulations are performed to predict thermal loads of the buildings. A simplified DGHCS is then designed and simulated to predict its energy use. A central borehole heat exchanger (BHE) of the DGHCS is sized using the RowWise algorithm of GHEDesigner to meet the thermal loads within user-specified land areas for installing BHE. The upgraded GeoWISE reports the needed capacity of heating and cooling equipment in each building, design of the central BHE, energy consumption reduction, and energy cost saving resulting from using DGHCS compared with conventional HVAC systems. A case study is showcased using the upgraded GeoWISE to design and conduct techno-economic analysis of a simplified DGHCS.

Prem Anand Jayaprabha, Jyothis Anand [ORNL] (ORCID↗

Superconducting Hyperdimensional Associative Memory Circuit for Scalable Machine Learning

Here we propose a generalized architecture for the first rapid-single-flux-quantum (RSFQ) associative memory circuit. The circuit employs hyperdimensional computing (HDC), a machine learning (ML) paradigm utilizing vectors with dimensionality in the thousands to represent information. HDC designs have small memory footprints, simple computations, and simple training algorithms compared to superconducting neural network accelerators (SNNAs), making them a better option for scalable SFQ machine learning (ML) solutions. The proposed superconducting HDC (SHDC) circuit uses entirely on-chip RSFQ memory which is tightly integrated with logic, operates at 33.3 GHz, is applicable to general ML tasks, and is manufacturable at practically useful scales given current SFQ fabrication limits. Tailored to a language recognition task, SHDC consists of ~ 2-20 M Josephson junctions (JJs) and consumes up to three times less power than an analogous CMOS HDC circuit while achieving 78-84% higher throughput. SHDC is capable of outperforming the state of the art RSFQ SNNA, SuperNPU, by 48-99% for all benchmark NN architectures tested while occupying up to 90% less area and consuming up to nine times less power. To the best of the authors' knowledge, SHDC is currently the only superconducting ML approach feasible at practically useful scales for real-world ML tasks and capable of online learning.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFit

Abstract Summary Bayesian inference in biological modeling commonly relies on Markov chain Monte Carlo (MCMC) sampling of a multidimensional and non-Gaussian posterior distribution that is not analytically tractable. Here, we present the implementation of a practical MCMC method in the open-source software package PyBioNetFit (PyBNF), which is designed to support parameterization of mathematical models for biological systems. The new MCMC method, am, incorporates an adaptive move proposal distribution. For warm starts, sampling can be initiated at a specified location in parameter space and with a multivariate Gaussian proposal distribution defined initially by a specified covariance matrix. Multiple chains can be generated in parallel using a computer cluster. We demonstrate that am can be used to successfully solve real-world Bayesian inference problems, including forecasting of new Coronavirus Disease 2019 case detection with Bayesian quantification of forecast uncertainty. Availability and implementation PyBNF version 1.1.9, the first stable release with am, is available at PyPI and can be installed using the pip package-management system on platforms that have a working installation of Python 3. PyBNF relies on libRoadRunner and BioNetGen for simulations (e.g. numerical integration of ordinary differential equations defined in SBML or BNGL files) and Dask.Distributed for task scheduling on Linux computer clusters. The Python source code can be freely downloaded/cloned from GitHub and used and modified under terms of the BSD-3 license (https://github.com/lanl/pybnf). Online documentation covering installation/usage is available (https://pybnf.readthedocs.io/en/latest/). A tutorial video is available on YouTube (https://www.youtube.com/watch?v=2aRqpqFOiS4&t=63s). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Adaptive Hierarchical Cyber Attack Detection and Localization in Active Distribution Systems

Development of a cyber security strategy for the active distribution systems is challenging due to the inclusion of distributed renewable energy generations. Here this paper proposes an adaptive hierarchical cyber attack detection and localization framework for distributed active distribution systems via analyzing electrical waveforms. Cyber attack detection is based on a sequential deep learning model, via which even minor cyber attacks can be identified. The two-stage cyber attack localization algorithm first estimates the cyber attack sub-region, and then localize the specified cyber attack within the estimated subregion. We propose a modified spectral clustering-based network partitioning method for the hierarchical cyber attack ‘coarse’ localization. Next, to further narrow down the cyber attack location, a normalized impact score based on waveform statistical metrics is proposed to obtain a ‘fine’ cyber attack location by characterizing different waveform properties. Finally, compared with classical and state-of-art methods, a comprehensive quantitative evaluation with two case studies shows promising estimation results of the proposed framework.

42 ENGINEERING↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A robust dynamic state estimation approach against model errors caused by load changes

Dynamic state estimation (DSE) plays an important role in power system security monitoring and online control. In practice, there are two approaches to implementing DSE. The first approach is distributed DSE, which is based on the assumption that the terminal bus of each generator can be measured by PMUs (phasor measurement units). The assumption cannot be satisfied currently, however, because PMUs usually are installed at important high-voltage buses such as 500-kV buses installed in portions of the grid overseen by the Western Electricity Coordinating Council. Another issue of this approach is that performance of DSE is vulnerable to bad measurement data. The reason for this vulnerability is that DSE is performed separately through measurements at each terminal bus, and measurements at terminal buses are the only measurement upon which DSE can rely. Therefore, important redundant measurements are not included in this approach. The second approach is centralized DSE. This approach does not have the requirement for PMU location, and redundant measurements can be considered fully. However, load changes and grid topology changes impact centralized DSE. In this paper, we propose a new approach for handling the impact of load changes on DSE. We have developed a new algorithm that includes two sequential steps. In the first step, errors caused by load changes are detected by analyzing the difference between prediction results and measured results. In the second step, once model error is detected, a model optimization procedure is run to correct the error so the state estimation error can be mitigated. Simulation results from the IEEE 68 bus system show that the proposed approach can effectively handle model errors caused by load changes.

robust dynamic state estimation, load change, powe↗

CoSHA: Code for Stellar Properties Heuristic Assignment—for the MaStar Stellar Library

We introduce CoSHA: a Code for Stellar properties Heuristic Assignment. In order to estimate the stellar properties, CoSHA implements a Gradient Tree Boosting algorithm to label each star across the parameter space (T eff , $\mathrm{log}g$, [Fe/H], and [α/Fe]). We use CoSHA to estimate the stellar atmospheric parameters of 22,000 unique stars in the MaNGA Stellar Library (MaStar). To quantify the reliability of our approach, we run internal tests, using both the Göttingen Stellar Library (a theoretical library) and the first data release of MaStar, and external tests, by comparing the resulting distributions in the parameter space with the APOGEE estimates of the same properties. In summary, our parameter estimates span the ranges T eff = [2900, 12,000] K, $\mathrm{log}g=[-0.5,5.6]$, [Fe/H] = [-3.74, 0.81], and αM = [-0.22, 1.17]. We report internal (external) uncertainties of the properties of ${\sigma }_{{T}_{\mathrm{eff}}}\sim 43(240)$ K, ${\sigma }_{\mathrm{log}g}\sim 0.2(0.4)$, σ [Fe/H] ~ 0.16(0.24), and σ [α/Fe] ~ 0.09(0.08). These uncertainties are comparable to those of other methods with similar objectives. Despite the fact that CoSHA is not aware of the spatial distributions of these physical properties in the Milky Way, we are able to recover the main trends known in the literature. The catalog of physical properties for MaStar can be accessed online.

79 ASTRONOMY AND ASTROPHYSICS↗

Multi-modal Energy-optimal Trip Scheduling in Real-time (METS-R) for Transportation Hubs (Final Report)

This report summarizes the work performed under the award number EE0008524. The project develops the Multi-modal Energy-optimal Trip Scheduling in Real-time (METS-R) platform as the next-generation transportation solution based on autonomous electric vehicles (AEV) serving passenger trips from and to urban transportation hubs, to substantially reduce transportation energy consumption. Extensive data collection and analyses were first conducted to understand the demand patterns and energy consumption of hub-based on-road trips. Then, a data-driven framework that consists of an analytical module and a simulation module was proposed. For the analytical module, five planning + operation tools were developed to support the planning and energy-efficient operations of urban AEV services: the charging station planning that robotically allocates charging supplies based on the stationary charging demand distribution; the transit planning and demand adaptive scheduling model that efficiently generates\ candidate transit routes from hubs to other places and dynamically adjusts the transit time table to fit the current demand; the online energy-efficient routing that learns the energy-optimal paths from observations of link-level energy consumption in real-time; the hub-based ridesharing that matches trip requests together with account for the uncertainty of future trip demand and vehicle supply; and finally, the integrated demand prediction and anomaly detection pipeline that leverages the flight/train time table and support other planning/operation tools. To demonstrate the performance of these tools, a scalable high-performance agent-based simulator was built. We divided the urban space into multiple service zones where each zone was considered as an agent for passenger generation and vehicle charging. Two types of AEV agents were coded to model two types of mobility services: AEV taxi and AEV transit. For the AEV taxi, the team implemented the functions of pickup/drop-off passengers, energy-efficient routing, ridesharing, fleet rebalancing, and recharging. For the AEV bus, the team implemented the functions of demand-adaptive route scheduling, passenger boarding, and recharging. A high-performance computing framework was introduced to receive various profiling information (such as link energy updates, vehicle speed) from the simulator instances and communicate the operational commands back to the instances. The numerical experiments show that each of the proposed operational algorithms can reduce energy consumption and improve system efficiency. Furthermore, there exists the need to collectively consider multiple planning + operational strategies as multiple strategies can influence each other in terms of performance impacts. Recommendations for future work related to AEV planning and simulation are discussed.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Classifying Seyfert Galaxies with Deep Learning

The traditional classification for a subclass of the Seyfert galaxies is visual inspection or using a quantity defined as a flux ratio between the Balmer line and forbidden line. One algorithm of deep learning is the convolution neural network (CNN), which has shown successful classification results. We build a one-dimensional CNN model to distinguish Seyfert 1.9 spectra from Seyfert 2 galaxies. We find that our model can recognize Seyfert 1.9 and Seyfert 2 spectra with an accuracy of over 80% and pick out an additional Seyfert 1.9 sample that was missed by visual inspection. We use the new Seyfert 1.9 sample to improve the performance of our model and obtain a 91% precision of Seyfert 1.9. These results indicate that our model can pick out Seyfert 1.9 spectra among Seyfert 2 spectra. We decompose the Hα emission line of our Seyfert 1.9 galaxies by fitting two Gaussian components and derive the line width and flux. We find that the velocity distribution of the broad Hα component of the new Seyfert 1.9 sample has an extending tail toward the higher end, and the luminosity of the new Seyfert 1.9 sample is slightly weaker than the original Seyfert 1.9 sample. This result indicates that our model can pick out the sources that have a relatively weak broad Hα component. In addition, we check the distributions of the host galaxy morphology of our Seyfert 1.9 samples and find that the distribution of the host galaxy morphology is dominated by a large bulge galaxy. In the end, we present an online catalog of 1297 Seyfert 1.9 galaxies with measurements of the Hα emission line.

79 ASTRONOMY AND ASTROPHYSICS↗

Dictionary Learning with Accumulator Neurons

The Locally Competitive Algorithm (LCA) uses local competition between non-spiking leaky integrator neurons to infer sparse representations, allowing for potentially real-time execution on massively parallel neuromorphic architectures such as Intel's Loihi processor. Here, we focus on the problem of inferring sparse representations from streaming video using dictionaries of spatiotemporal features optimized in an unsupervised manner for sparse reconstruction. Non-spiking LCA has previously been used to achieve unsupervised learning of spatiotemporal dictionaries composed of convolutional kernels from raw, unlabeled video. We demonstrate how unsupervised dictionary learning with spiking LCA (\hbox{S-LCA}) can be efficiently implemented using accumulator neurons, which combine a conventional leaky-integrate-and-fire (\hbox{LIF}) spike generator with an additional state variable that is used to minimize the difference between the integrated input and the spiking output. We demonstrate dictionary learning across a wide range of dynamical regimes, from graded to intermittent spiking, for inferring sparse representations of both static images drawn from the CIFAR database as well as video frames captured from a DVS camera. On a classification task that requires identification of the suite from a deck of cards being rapidly flipped through as viewed by a DVS camera, we find essentially no degradation in performance as the LCA model used to infer sparse spatiotemporal representations migrates from graded to spiking. We conclude that accumulator neurons are likely to provide a powerful enabling component of future neuromorphic hardware for implementing online unsupervised learning of spatiotemporal dictionaries optimized for sparse reconstruction of streaming video from event based DVS cameras.

artificial intelligence↗

Effects of overlapping sources on cosmic shear estimation: Statistical sensitivity and pixel-noise bias

The next generation of dark-energy imaging surveys — so called “Stage-IV” surveys, such as that of the Rubin Observatory Legacy Survey of Space and Time (LSST) — will cross a threshold in the number density of detected sources on the sky that requires qualitatively different image analysis and measurement techniques compared to the current generation of Stage-III surveys. In Stage-IV surveys, a significant amount of the cosmologically useful information is due to sources whose images overlap with those of other sources on the sky. Here, we focus on the weak gravitational lensing probe, for which we expect the largest impact since the cosmic shear signal is primarily encoded in the estimated shapes of observed galaxies and thus directly impacted by overlaps. We introduce a framework based on the Fisher formalism to analyze the effect of the overlapping sources (“blending”) on the estimation of cosmic shear. This method gives concrete predictions for the minimum loss of information due to noise and blending for any choice of “deblending” scheme and shape-measurement algorithm. Our studies account for undetected sources but do not address their full effects and biases they may introduce. We use simulated images and predict this impact of blending for three surveys: the Dark Energy Survey (DES), the Hyper-Suprime Cam Subaru Strategic Program (HSC-SSP), and the Rubin LSST. Our methodology successfully estimates the statistical sensitivity to weak lensing for DES and HSC early results. For LSST, we present the expected loss in statistical sensitivity for the ten-year survey due to blending. We find that for approximately 62% of galaxies that are likely to be detected in full-depth LSST images, at least 1% of the flux in their pixels is from overlapping sources. We also find that the statistical correlations between measures of overlapping galaxies and, to a much lesser extent (0.2%) the higher shot noise level due to their presence, decrease the effective number density of galaxies, N eff , by ~ 18%. We calculate an upper limit on N eff of 39.4 galaxies per arcmin 2 in r band. We study the impact of stars on as a function of stellar density and illustrate the diminishing returns of extending the survey into lower Galactic latitudes. We extend the simulation-based Fisher formalism to predict the expected increase in pixel-noise bias due to blending for maximum-likelihood (ML) shape estimators. We find that noise bias depends sensitively on the particular shape estimator and measure of ensemble-average shape that is used, and properties of the galaxy that include redshift-dependent quantities such as size and luminosity. The source code for these studies is available online.[The documented software developed for the catalog-level studies are available in the open-source LSST DESC github repository https://github.com/LSSTDESC/WeakLensingDeblending. The software for analyzing one or two galaxies with user-defined parameters is in the open-source github repository https://github.com/ismael-mendoza/ShapeMeasurementFisherFormalism.]

79 ASTRONOMY AND ASTROPHYSICS↗

Data-driven linear time advance operators for the acceleration of plasma physics simulation

In this study, we demonstrate the application of data-driven linear operator construction for time advance with a goal of accelerating plasma physics simulation. We apply dynamic mode decomposition (DMD) to data produced by the nonlinear SOLPS-ITER (Scrape-off Layer Plasma Simulator - International Thermonuclear Experimental Reactor) plasma boundary code suite in order to estimate a series of linear operators and monitor their predictive accuracy via online error analysis. We find that this approach defines when these dynamics can be represented by a sequence of approximate linear operators and is essential for providing consistent projections when compared to an unconstrained application. For linear diffusion and advection–diffusion fluid test problems, we construct and apply operators within explicit and implicit time advance schemes, demonstrating that stability can be robustly guaranteed in each case. We further investigate the use of the linear time advance operators within several integration methods including forward Euler, backward Euler, and the matrix exponential. The application of this method to simulation data from SOLPS-ITER, with varying levels of Markov chain Monte Carlo numerical noise, shows that constrained DMD operators yield a capability to identify, extract, and integrate a (slow) subset of the present timescales. Example applications show that for projected speedup factors of [Formula: see text], and [Formula: see text], a mean relative error of 3%, 5%, and 8% and maximum relative error less than 20% are achievable, which appears acceptable for typical SOLPS-ITER steady-state simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗