Leveraging Pre-trained Neural Networks for Image Classification in Audio Signal Analysis for Mobile Applications of Home Automation
V. I. Slyusar, I. I. Sliusar · River Publishers eBooks · 2024
This chapter presents an in-depth analysis of innovative approaches in the field of audio signal classification using convolutional neural networks (CNNs) and their integration with image processing techniques. We investigate the effectiveness of transfer learning from image to audio domains, examining various neural network architectures like VGG16, DenseNet201, MobileNetV3Small, and EfficencyNet. Special emphasis is placed on the adaptability of these networks to handle audio data, particularly through the manipulation of input sizes and structures, such as Mel-frequency cepstral coefficients (MFCCs) and short-time Fourier transform (STFT) spectrograms. Significant findings include the discovery that pre-trained image classification networks can be effectively repurposed for audio signal analysis. By adjusting parameters such as learning rate and batch size, and experimenting 110 with different architectures, we achieved considerable improvements in classification accuracy. For instance, replacing MobileNet with a pre-trained EfficencyNet in a specific architecture resulted in a record accuracy of 82% at the 134th epoch. Additionally, we explore the potential of these architectures for the unified processing of audio and image data, suggesting a method for task-specific commutation of input signals and weight loading in non-pre-trained layers. This approach highlights the versatility and potential of neural networks in handling diverse data types beyond their initial training scope.