Interpreting Deep Neural Networks via Relative Activation-Deactivation Abstractions
Zhen Zhang, Peng Wu, Yuting Yang, Xuran Li · ACM Transactions on Software Engineering and Methodology · 2025
As deep learning models are widely applied in various real-world intelligent systems, the interpretability and trustworthiness of these models have attracted substantial attention. A succinct and effective abstraction that can represent the inference behavior of a deep neural network is significant for explaining its decision logic and ensuring its reliability. We propose in this paper the relative activation and deactivation patterns to redefine the behaviors of deep learning neurons, and the notion of relative selectivity to quantify the output differences of neurons among different prediction categories. Then, we present a relative activation-deactivation abstraction approach to characterize the decision logic of a deep learning model. The relative activation-deactivation abstractions enjoy close intra-class aggregation for each prediction category, as well as diverse inter-class separation between categories. This abstraction approach can be well extended from CNNs to Transformers, demonstrating its excellent scalability. We further propose an anomaly detection algorithm based on the relative activation-deactivation abstraction approach, following the principle that the relative activation-deactivation abstraction of a deep learning model under an abnormal input is far away from the one for the predicted category the deep learning model outputs. We evaluate the anomaly detection algorithm with 13 typical benchmark datasets in the computer vision and natural language processing domains. The experimental results on two widely concerned anomaly detection tasks, i.e., out-of-distribution and adversarial detection, show that our algorithm can achieve higher and more stable detection performance than the state-of-the-art algorithms, with significantly more true positives and fewer false positives for anomaly detection.