Trojan Attack and Defense for Deep Learning Based Power Quality Disturbances Classification

Sultan Uddin Khan, Mahmoud Nabil, Mohamed M. E. A. Mahmoud, Maazen Alsabaan, Tariq A. Alshawi · IEEE Transactions on Network Science and Engineering · 2025

Accurate classification of power quality disturbances (PQD) is essential for ensuring the reliability and safety of modern power systems. However, deep learning (DL) models used for PQD classification can be compromised by trojan attacks—malicious modifications that alter model behavior only in the presence of specific triggers. Motivated by the urgent need to safeguard modern power grids, we, for the first time, propose trojan attack for DL-based PQD classification by introducing a novel trojan attack algorithm called Sneaky Spectral Strike ($S^{3}$). Key features of$S^{3}$include balancing the signal-to-noise (SNR) ratio to maintain imperceptibility and optimizing the fooling rate (FR) for maximum effectiveness. Leveraging the Fast Fourier Transform (FFT) and trigger optimization techniques to embed a stealthy trigger,$S^{3}$achieves an impressive fooling rate of 99.9% with minimal impact on clean-data accuracy across diverse DL architectures. Unlike prior time series data (TSD) based trojan attack studies that use datasets with limited sample diversity and time-step variations, we utilize a comprehensive PQD dataset encompassing a wide range of events and varied time steps, thereby exposing vulnerabilities in more diverse and realistic scenarios. S3outperforms state-of-the-art methods, improving the average fooling rate by 7.4% over TimeTrojanDE, 0.83% over TSBA, and 0.25% over TrojanFlow. To assess generalizability, S3was also evaluated on two additional time-series datasets, achieving fooling rate of 99.89% on the Online Retail Dataset and a perfect 100% on the Household Electric Power Consumption Dataset. To counter such sophisticated attacks, we also propose an innovative defense mechanism that detects trojan attacks by analyzing decision boundary discrepancies resulting from trojan insertion and injecting universal adversarial perturbations. Our defense strategy demonstrates effectiveness in identifying compromised models across various class scenarios, with the capability to detect infected models across 14 of the 17 class scenarios with trojan infection probability peaks at 0.94076.

Read the paper · More papers on PaperTik