Interpretable intrusion detection for IoT environments using a self-attention-based explainable AI framework
Kanta Prasad Sharma, Tapsi Nagpal, Tarak P. Vora, Anupam Yadav, Muhammad Irsyad Abdullah, B. Jayaprakash, Aditya M Kashyap, G. Sridevi, Ayan Bhowmik, Bethelehem Burju Bukate · Scientific Reports · 2025
The rapid growth of Internet of Things (IoT) devices has introduced critical cybersecurity challenges due to the large attack surface and diverse interconnected systems. This work proposes a novel self-attention deep neural network (SA-DNN) augmented with a learnable feature gating (LFG) mechanism. The LFG component dynamically emphasizes security-relevant features during training, while suppressing redundant data. This architecture enables seamless, end-to-end feature optimization. It enhances the robustness of the model across varied intrusion scenarios without the need for conventional manual or filter-based feature selection. The self-attention module efficiently captures long-range relationships within network traffic, while the LFG mechanism boosts the adaptability of feature representations. The framework is explainable through shapley additive explanations (SHAP) and local interpretable model-agnostic explanations (LIME) and provides transparency in intrusion detection decisions. Evaluations on IoT datasets BoT-IoT and N-BaIoT show superior performance of the proposed method over baseline and state-of-the-art methods, achieving accuracy values of 99.3% and 99.6%, respectively. Additional experiments on the general network dataset UNSW-NB15 demonstrate strong generalizability with 97.9% accuracy. This confirms the robustness of the model across diverse intrusion scenarios. The proposed approach provides a lightweight, scalable, and interpretable solution with strong generalizability across IoT and broader network security settings.