CNN based Interpretability Analysis of Voiceprint Recognition
Zhehan Wu, Xiulan Hao · Zenodo (CERN European Organization for Nuclear Research) · 2023
Voiceprint recognition is a type of biometric recognition which converts acoustic signals into electrical signals and then uses computers for recognition. Voiceprint recognition is widely used, for example, power companies use it to detect equipment failures, while environmental protection administration institutions use it to monitor ecological status. In particular, when used in bird sound recognition, we can monitor ecological diversity of jungles, swamps, to prevent birds from colliding with aircraft, and to provide a rigor reference for the differential prevention and control of electric transmission line bird failures. The community have devoted to various models that use deep learning to improve voiceprint recognition. However, the popular neural network models suffer from insufficient interpretability, so do neural network-based voiceprint recognition models. In order to interpret the model behavior more intuitively, such as how deep learning makes such a prediction, why some features are favored over others by a model, we look for the determinants of the spectrogram for the final voiceprint recognition in an image recognition network through an interpretable analysis of the neural network. We use a CNN-based image recognition network and apply it to the more challenging Kaggle Rainforest Bird dataset. In this interpretable framework, it is modified such that filters are encouraged to learn more interpretable parameters so that features can be easily identified in a spectrogram. Spectrograms where bird sound occurs are called key spectrograms, while spectrograms where no bird sound occurs are called non-key spectrograms. Experimental results visualize the differences of key spectrograms from non-key spectrograms in the time and frequency dimensions. Meanwhile, we compare the key spectrograms of higher recognition rate with those of lower recognition rate and find the former have more salient features