Exploring Transfer Learning Approach for Environmental Sound Classification: A Comparative Analysis

Theresia Herlina Rochadiani, Yulyani Arifin, Derwin Suhartono, Widodo Budiharto · 2024

The expansion of deep learning techniques, as well as the availability of large audio/sound datasets, have fueled tremendous breakthroughs in audio/sound classification during the last several years. The transfer learning approach has emerged as one of the primary approaches for improving the accuracy and durability of classification systems. This study conducts a comprehensive comparative analysis to determine the effectiveness and performance of this method in environment sound classification. This current investigation focuses on environmental sound classification using VGGish and YAMNet pre-trained models with the ESC-50 and BDLib2 datasets. In the ESC-50 dataset, VGGish improves accuracy to 372.22%, while YAMNet improves accuracy to 383.33% when compared to baseline models. Similarly, in the BDLib2 dataset, accuracy increases significantly to 221.43% with VGGish and 246.43% with YAMNet. Transfer learning exhibits remarkable effectiveness in enhancing model performance, with significant accuracy boosts observed in both datasets. YAMNet, designed specifically for sound classification tasks, surpasses VGGish in improving environmental sound classification performance, potentially due to its architecture’s adaptability and diverse training on environmental sounds.

Read the paper · More papers on PaperTik