Adversarial Examples Detection Using Random Projections of Residuals

Kaijun Wu, Anjie Peng, Yao Wu Hao, Qi Zhao · 2020

The adversarial images make the deep neural network misclassify the original label and thus fool successive artificial intelligence system. The emergency of the adversarial images attracts researchers' attentions to the security of machine learning. In this paper, we employ a steganalysis based method to detect adversarial images which are generated by the typical attacks including BIM and DEEPFOOL. Unlike the previous steganalysis-based methods, we project the residuals onto randomly neighborhood and extract the histogram as the feature. Compared with steganalysis-based methods which calculates the co-occurrence on the truncated residual as feature, the proposed method does not truncate the residual thus avoiding loss information. Experimental results show that the proposed method obtained better performance than SPAM and SRM based methods for BIM and DEEPFOOL attack.

Read the paper · More papers on PaperTik