Combinatorial Evaluation of Physical Feature Engineering, Classical Machine Learning, and Deep Learning Models for Synchrophasor Data at Scale

Sean Patrick Murphy, Mohini Bariya, Debbie Chang, Jeff P. Lin, Chris Ryan, Ramiro Mata · 2022

A major objective of the project was to train and evaluate the effectiveness of multiple event and anomaly detection, identification and classification deep temporal learning models for processing of real-time phasor measurement unit (PMU) data streams. A vast dataset, consisting of two years of phasor measurements from all three U.S. Interconnections, was curated and released by the Department of Energy (DOE) through Pacific Northwest National Laboratory (PNNL). The dataset also included an event log that provided event times and types (e.g. generator trips, line trips, planned service events, transformer operations, etc.). Our analysis of this dataset addressed six (6) of the eleven (11) research priorities identified in Funding Opportunity Announcement (FOA) DE-FOA-0001861 “Big Data Analysis of Synchrophasor Data” (FOA 1861). Rather than being limited to pre-determined specific algorithms, this project relied on the uniquely structured, highly performant underlying time series database capabilities of the PredictiveGrid platform to assess the vast dataset utilizing a wide variety of algorithms. Event detection and classification was the primary focus of this project; specifically developing an understanding the characteristics and types of events recorded in the event log as well as—what we believed to be—events appearing in the synchrophasor data but not identified in the event log. Spatial and temporal patterns in events were also investigated but the absence of any topology information in this dataset made such analysis challenging. After extensive data exploration, we concluded that events manifest as anomalies according to some statistical metric. Therefore, to detect events, we created a set of statistical event detectors. The detectors were run on voltage magnitude streams because these particular measurements mostly hover around a nominal value and therefore lend themselves to statistical anomaly detection. Furthermore, PredictiveGrid precomputes and stores aggregate statistics for various temporal data resolutions and thus our statistical event searches were able to rapidly run across months of high resolution data.. Numerous statistically significant events were detected that were visible in the measurement streams of several phasor measurement units (PMUs) but absent from the utility event logs. We found that most utility logged events also manifest as statistical anomalies in the measurement streams, buttressing the hypothesis that operational significance generally implies statistical significance (the converse is not verifiable in the absence of additional expert feedback). Alongside the statistical event detectors—which seek to discover events independently of any logs—we developed a heuristic approach, termed Heuristic Event Localization (H-loc), to refine the event information in the logs themselves. Starting from the event time recorded in the log, H-loc searches for large deviations across sensors and streams. The sensor and its stream with the largest deviation are returned as the event epicenter and manifesting data type respectively, and the precise time of the deviation is returned as the refined event time. We found that this preprocessing step improves the performance of downstream analytics enormously, as it extricates the event signal from the overpowering background “noise” which results from inaccurate timestamps and the absence of spatial specificity. We also applied machine learning models to the tasks of event detection and classification. We trained and tested a suite of models with varied architectures and input feature types. By individually applying several simpler models with single feature types as inputs, rather than taking an ensemble approach, we aimed, not simply to maximize performance, but to reveal which feature types and model types were best suited to this relatively novel context. This we believed would be the more valuable finding to guide future work. All machine learning (ML) models ingested a window of data across multiple stream types from a single sensor. The models were trained on four tasks: determining if the data window contained an event; determining if the data window contained an event precursor (i.e. if an event was to follow in a fixed time frame), for a data window known to contain an event, determining the event class out of a set of classes; for a data window known to contain an event precursor, determining the event class of the impending event out of a set of classes. The first two of these tasks are binary, while the latter two are multiclass. We achieved good performance on the event detection task as well as—perhaps surprisingly—on the precursor detection task. Performance on the classification tasks was low across all models. This suggests the presence of a relatively strong discriminating signal during or just before an event. However, identifying the type of the event is more challenging, probably requiring more labeled event examples and further feature engineering. The strongest message we take from this work is the need for tooling to determine which parts of the data are consequential and require attention. Discovery is an enormous challenge in large data volumes; if interesting periods are not localized, both spatially and temporally, they are lost in the sheer quantity of data and the continual baseline variation of grid measurements. This localization is critical for human users and algorithmic users, as we found when comparing analytic performance on the raw event logs and those refined by H-loc. Determining which periods are important requires a combination of domain knowledge and a knack for computational optimization, the specifics of which will depend on the data platform. Optimization is critical as analyses must be run across all the data, not just on narrow regions around logged event times. For example, we discovered many seemingly significant events outside the event logs. It appears that there is much going on in the system that, at best, utilities have simply not recorded in their logs, and, at worst, have no awareness of.

Read the paper · More papers on PaperTik