Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “synchrophasor datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Using Synchrophasor Status Word as Data Quality Indicator: What to Expect in the Field?

Data quality plays a crucial role in successful applications of synchrophasor data in power system operation and control. This paper presents the results of a data quality analysis of a multi-year field-recorded synchrophasor dataset. The analysis has identified several typical data quality issues encountered in the field data. An examination of the PMU status words included with the dataset has revealed several inconsistent implementations and the lack of correlation between the PMU data quality and the status word, which impacts the usefulness of such information. Our investigation has concluded that the status word alone as found in the recorded field dataset could not be used as a reliable indicator of data quality for field-recorded data. Several recommendations are proposed to improve the usefulness of the PMU status word.

Cheng, Zheyuan↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evaluating a Commercial Dynamic Line Rating Software with the National PMU Dataset

To accelerate the development of data-driven applications for power systems, the Department of Energy (DOE) supported the collection and curation of a synchrophasor dataset spanning two years of observations from transmission utilities across the US. This National PMU Dataset (NPDS) was anonymized and distributed to awardees of a DOE research grant under nondisclosure agreements (NDAs) but has also been retained at PNNL to enable further research. Agreements with data contributors prevent the data from being shared outside the organization. However, establishing a blind research validation methodology is envisioned to maximize the value proposition of the NPDS. In this validation strategy, researchers may share algorithms/software (potentially as executables to protect intellectual property) with PNNL, and PNNL will share feedback about the software’s performance on subsets of the NPDS. Such a blind methodology ensures that sensitive information about critical infrastructure remains protected, but the value of the NPDS can be extended to research beyond PNNL. Through iterative feedback, the algorithms may be tweaked to address real-world artifacts. As the NPDS data is temporally and geographically diverse, it may capture features absent in smaller datasets used during the development of the algorithm under test. This report presents lessons learned from applying the blind validation methodology to LineID™, a synchrophasor-based dynamic line rating software developed by Topolonet Corporation. Improvements made to the software through iterative feedback, limitations of the validation methodology, as well as how the limitations of the NPDS affected the evaluation process are discussed. Observations indicate that the proposed validation methodology can be valuable for evaluating other tools in the future.

97 MATHEMATICS AND COMPUTING↗

Open Source Synergy: Developing and Validating PMU Data Analysis Techniques Using Open Source Tools and Datasets

This paper presents an exploration into the development and validation of data analysis approaches for Phasor Measurement Units (PMUs) using open-source datasets and tools. Various methods for event detection, event classification, frequency response, and oscillation analysis were tested. We leverage the capabilities of Archive Walker (AW), the Frequency Response Analysis Tool (FRAT), and the Oscillation Baselining and Analysis Tool (OBAT), all open-source tools, for efficient processing and analysis of synchrophasor data. The open-source Transmission Signature Library (TSL) dataset was employed as a dataset for a comprehensive evaluation to assess the performance and reliability of the proposed methods.

PMU, event analysis, oscillation, Frequency Respon↗

Combinatorial Evaluation of Physical Feature Engineering, Classical Machine Learning, and Deep Learning Models for Synchrophasor Data at Scale

A major objective of the project was to train and evaluate the effectiveness of multiple event and anomaly detection, identification and classification deep temporal learning models for processing of real-time phasor measurement unit (PMU) data streams. A vast dataset, consisting of two years of phasor measurements from all three U.S. Interconnections, was curated and released by the Department of Energy (DOE) through Pacific Northwest National Laboratory (PNNL). The dataset also included an event log that provided event times and types (e.g. generator trips, line trips, planned service events, transformer operations, etc.). Our analysis of this dataset addressed six (6) of the eleven (11) research priorities identified in Funding Opportunity Announcement (FOA) DE-FOA-0001861 “Big Data Analysis of Synchrophasor Data” (FOA 1861). Rather than being limited to pre-determined specific algorithms, this project relied on the uniquely structured, highly performant underlying time series database capabilities of the PredictiveGrid platform to assess the vast dataset utilizing a wide variety of algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Open-Access Repository of Synchrophasor Data Quality Examples: Curation and Example Applications

Synchrophasor measurements are critical in providing wide-area situational awareness to power system operators. However, data artifacts may be introduced due to various issues such as loss of communication, loss of GPS signal, internal clock error, and vendor-specific implementation of phasor estimation algorithms. Tools designed to provide actionable insights from synchrophasor data, hence, must be designed to be robust to these data quality issues. In this work, two years of synchrophasor data sourced from multiple electric utilities in the United States were analyzed to identify examples of data quality problems. These examples were then labeled and published in the Grid Event Signature Library, a publicly available repository of power system measurements hosted by the Oak Ridge National Laboratory. This paper describes the data curation process, and illustrates two application use cases where the dataset can be valuable to the research community. In the first use case, a random forest classifier is trained to distinguish power system disturbance signatures from data anomalies introduced in synchrophasor measurements due to clock errors. The second use case studies the impact of data quality issues on an example synchrophasor application (specifically, event start time determination). The choice of data quality problems investigated is informed by the examples in the repository curated in this work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning Guided Operational Intelligence from Synchrophasors (Final Report)

Schweitzer Engineering Laboratories (SEL) and Oregon State University (OSU) received over 27 terabytes of electrical power system phasor measurement unit (PMU) data for the Eastern, Western, and ERCOT interconnections. The dataset includes measurements spread across 446 PMUs from early 2016 to mid 2018 depending on the interconnect. The full dataset was split into a training and test (holdout) dataset by PNNL. All data was received in the Apache Parquet format. The overarching goal of this project is to develop and execute a strategy to mitigate data anomalies, perform analysis on the dataset, and detect anomalous events in the data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FOA 1861 Data Curation Overview

This document describes the process executed to collect, examine, and consolidate Phasor Measurement Unit (PMU) data from multiple transmission operators into a common dataset. The consolidated PMU data set was further anonymized and distributed to the Department of Energy Funding Opportunity Announcement (FOA) 1861 Big Data Analysis of Synchrophasor Data awardees.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Use of Machine Learning on PMU Data for Transmission System Fault Analysis

Synchrophasor technology has been used for monitoring, control, and protection of bulk power system for over 10 years. Deployment of phasor measurement units (PMUs) in the USA power system has surpassed 3000 units installed in the transmission substations as stand-alone intelligent electronic devices (IEDs) or as a software add-on to other devices such as digital protective relays (DPRs) or digital fault recorders (DFRs). By now, thousands of terabytes of PMU data may have been captured and stored by various transmission system operators (TSOs) and independent system operators (ISOs). This creates an opportunity to deploy advanced machine learning (ML) techniques to detect and classify faults recorded by PMUs automatically to be used by the system operators for rapid, critical decision-making when manual analysis of the past or unfolding events is not feasible. In this paper we offer a brief background on how the automated fault analysis may be done using DPR and/or DFR data, and compare some of the legacy approaches to the new ML approaches in the context of the system-wide PMU recordings. We then offer insights from developing practical ML solutions that have been applied on field recordings captured by close to 450 PMUs from all three US interconnections (Western, Eastern and ERCOT) over two years (2016-2017). We identify and illustrate ML challenges we addressed: inaccurate data, data with scarce and temporally imprecise fault labels, data recorded by PMUs sparsely located at substations resulting in the fault records taken afar from the ends of the faulted lines, data containing only positive sequence values, and data taken at different voltage levels. We then illustrate the ML model results for fault analysis under different application scenarios. The novelty of this study is not only in the design, implementation, and performance analysis of the ML algorithms, but also in the use of advanced fault modelling and simulation approaches to improve the training results when developing supervised ML models for fault detection and classification. Extensive simulations of faults were conducted on a 14-bus power system to create a training dataset with over 1400 accurately labelled faults. This dataset was applied to enhance the accuracy of fault detection and classification of machine learning-based models trained with small number of labelled faults in large datasets recorded in the grid interconnections ranging from 5,000 to 70,000 buses.

Synchrophasors, Machine Learning, Fault Analysis, ↗

Bayesian High-Rank Hankel Matrix Completion for Nonlinear Synchrophasor Data Recovery

Phasor measurement units (PMUs) provide high temporal-resolution synchrophasor measurements for power system monitoring and control. The frequent data quality issues, such as missing and bad data, prevent the incorporation of synchrophasor data in real-time operations. Most existing data-driven data recovery methods assume the power system dynamics can be approximated by a linear dynamical system, and the recovery performance degrades significantly when the power system is experiencing nonlinear dynamics during significant events. Here, this paper proposes a data-driven Bayesian nonlinear synchrophasor data recovery method (Ba-NSDR) that can recover a consecutive time period of simultaneous data losses or errors across all channels, even when the underlying system is highly nonlinear. The idea is to lift the Hankel matrix of the spatial-temporal synchrophasor data to a higher dimension such that the lifted Hankel matrix is low-rank in that space and can be processed with the kernel trick. Our proposed Bayesian method then infers the probabilistic distributions of synchrophasor from the partial observations. Some distinctive features of Ba-NSDR include an uncertainty index to measure the accuracy of the recovery result and the robustness to parameter selections. Our method is verified on both synthetic and recorded event datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Anomaly Detection, Localization and Classification using Drifting Synchrophasor Data Streams

With ongoing automation and digitization of the electric power system, several Phasor Measurement Units(PMUs) have been deployed for monitoring and control. PMU data can have multiple anomalies, and many of the researchers in the past have concentrated on training machine/deep learning algorithms offline for anomaly detection over PMU data (i.e., not in real time). These machine/deep learning algorithms, when trained offline on a sample rather than a population of the dataset, fail to consider the dynamic behavior of the power grid in real-time, resulting in low accuracy. Considering the dynamic behavior of the power grid (e.g., change in load, generation, distributed energy resources (DERs) switching, network, controls), the definition of data anomalies varies in time and requires online training. A fundamental challenge is to enable online (i.e., real-time) training of machine/deep learning algorithms for anomaly detection over streaming PMU data. While machine/deep learning is often desirable to manage data streams, training a deep learning algorithm over streaming PMU data is nontrivial due to changes in data statistics caused by dynamic streaming data. This paper proposes PMUNET: a novel device-level deep learning-based data-driven approach for anomaly detection, localization, and classification over streaming PMU data, using online learning and multivariate data-drift detection algorithm .Two variants of PMUNET, Dynamic data Change Driven Learning (DCDL) and Continuity Driven Learning (CDL), are proposed and compared. DCDL aims to train the deep learning algorithm whenever the definition of anomaly changes due to the power grid dynamics. On the other hand, CDL continuously trains the deep learning algorithm over the PMU data-stream. The experimental results verify that DCDL outperforms CDL and other efficient anomaly detection methods over multiple events such as faults and load/ generator/capacitor/DERs variations/switching for IEEE 14 and 39 Bus test system as well as real PMU industrial data. The result verifies that DCDL variant of PMUNET improves over existing approach with a gain of 2% - 10% in terms of accuracy, false-positive rate, and false-negative rate.

adversarial deep learning↗

Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART)

This report contains key findings from a project titled Big Data Synchrophasor Monitoring and Analytics for Resiliency Tracking (BDSMART), which was carried out through a collaborative effort of a team of researchers from Texas A&M Engineering Experiment Station, Temple University, and Quanta Technology, LLC. The in-kind support came from OSIsoft (acquired by AVEVA), which provided their PI Historian software to demonstrate the use case of streaming PMU data. The first section of the report describes the project goals and objectives related to the development of Machine Learning (ML) models capable of detecting and classifying events by processing phasor measurements captured in the field by Phasor Measurement Units (PMUs). The data for this study was contributed by the utilities/ISOs from the Western and Eastern interconnects and ERCOT, further referred to as Interconnect B (IC B), Interconnect A (IC A), and Interconnect C (IC C), respectively. The approach that the BDSMART Research Team proposed and the key research tasks defined by the team are outlined in this section. The next section describes the technical approach. We first discuss the data constraints related to the PMU measurements and data interpretation constraints imposed by the data contributors. They provided neither the topological information of the grid nor PMU placement locations and captured recorded data at very few locations in the system with the reporting rate of either 30 or 60 fps. The recordings are mostly positive sequence voltage, frequency, and ROCOF, and in some limited cases, three-phase voltages and currents. We then reflect on the bad data issues that stem from poor recording practices and vague definitions of the PMU status bits to supposedly be used for bad data identification. Finally, the data discovery points to imprecise time stamps with incomplete event start/end time, as well as inconsistent and incomplete event labeling, which combined make the implementation of the data models using supervising learning quite challenging. Following the data discovery study, we hypothesize that because the IC B data has the most complete labels, we should focus our model development on that data and then test it on data from other interconnects. We also define the common metrics used to evaluate the results from the ML algorithm tests. We concluded this section by summarizing the common ML models we used and explaining how we implemented and tested them. The issues from this section are expanded in the Training Dataset Report from this project. The final section of this report deals with the accomplishments and conclusions. As the accomplishments, we formulate the problem we are solving and what is achieved by solving the problem. We then reflect on each of the analytics tools we developed and point out the performance of each tool when applied to solving the mentioned problems. We reference this work for further details to the papers we published on each tool. In the conclusions, we give recommendations on how to improve future PMU recording practices to facilitate the ML algorithm implementation and guidance for the future standardization work aimed at clarifying the ambiguities associated with the PMU status bits. We finally list future tasks that can bring about further improvements in the proposed algorithms. The issues from this section are expanded in the Training, and Test Dataset Report filed at the project completion date.

97 MATHEMATICS AND COMPUTING↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

A National Infrastructure for Artificial Intelligence on the Grid (NI4AI) (Final Scientific/Technical Report)

Electric utilities have traditionally taken a very pragmatic yet myopic approach with grid sensors and the resulting collected data. Sensors are purchased and deployed to solve a specific, known problem that has risen to sufficient awareness as to justify the effort of deploying sensors and the needed capital investment. This sensor data flows into proprietary software packages with limited functionality intended only to address the initial problem. This approach aligns with the financial incentives of the utility to deploy capital into fixed hardware assets for which the corporations earn a rate of return. This mentality stands in stark contrast to the big data revolution that started nearly 25 years ago with the rise of Google. In this worldview, data is a fundamental business asset; successful organizations collect, store, explore, merge, and exploit as much data as possible to not only solve problems well understood today but also to tackle new problems that will inevitably rise tomorrow. The ARPA-E Open Innovation 2018 project entitled A National Infrastructure for Artificial Intelligence on the Grid or NI4AI for short was designed to demonstrate this alternative paradigm for using data. To do this, the project was composed of three key thrust areas. The first major component deployed a variety of high-frequency grid sensors and captured terabytes of both wide-scale and localized grid measurements, generating high-value datasets for grid research and algorithm development. The second aspect made available PingThings’ PredictiveGridTM, a horizontally scalable, cloud-based data management and AI platform built for time series data to explore and exploit the collected data. Finally, the project fostered a diverse and open research community composed of experts from numerous fields through focused educational content, code sharing, and data science competitions. Shifting away from “single use” sensors and closed data silos within electric utilities is a major benefit to the public at large. This legacy approach to data is incredibly (1) capital intensive (new sensors must be deployed for each new problem and problems tend to arise continuously) and (2) painfully slow (new problems must be identified first and then new sensors must be deployed to collect data to begin to address the issue). The transition to a carbon neutral grid requires a massive transformation of the existing grid infrastructure and will continue to challenge the legacy grid in unforeseen ways. The only way to make the energy transition cost effective is for utilities to abandon this dated data paradigm and adopt more contemporary approaches. NI4AI has shown that it is technically possible and economically feasible to ingest, explore, and exploit grid data collected from even very high frequency sensing, such as continuous point on wave sensors collecting measurements 10,000 times a second. In fact, the PredictiveGrid platform used is commercially available and deployed at several utilities in the United States. Project accomplishments were numerous and included (1) making available a state of the art time series platform to the community, (2) collecting over 520 streams of time series data from grid sensors totaling over 1 trillion grid measurements, and (3) developing and nurturing a community within the industry focused on the use of data to create value for utilities and, ultimately, end consumers.

97 MATHEMATICS AND COMPUTING↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗