Revamped Knowledge Distillation for Sound Classification
Achyut Mani Tripathi, Aakansha Mishra · 2022 International Joint Conference on Neural Networks (IJCNN) · 2022
This paper presents a novel knowledge distillation technique that inherits knowledge from multiple deep Environment Sound Classification (ESC) models trained on spectrogram features created by dividing the spectrogram into multiple subband spectrogram. The deep models trained on sub-band spectrograms prevent information loss while performing knowledge distillation from a teacher model to a student model receiving the full spectrogram as an input. The student models' performance is evaluated on two benchmark sound datasets, viz. the ESC-10 and Audio MNIST datasets. The impact of teacher models trained with different number of sub-band features and four ensemble techniques has been investigated thoroughly to enhance the final accuracy of the student model supervised by the proposed knowledge distillation framework. Experiments and results shows that the accuracy of the student model is comparable and competitive to state-of-the-art methods for sound classification. Moreover, the student model trained on the Audio MNIST dataset attains an hitherto unpublished accuracy of 98.25%, a new benchmark for the Audio MNIST dataset. Additionally, Grad- CAM visualization of the spectrogram features is generated to identify the spectrogram's relevant regions and understand why the model classifies a signal into a specific class.