Optimized Audio Classification with Convolutional Neural Network Ensembles
K V Sudheesh, Kalavala Swetha, C N Gireesh Babu, H N Ravikiran, Kiran Puttegowda, D S Sunil Kumar · 2024
In this work, we present a structured approach to a typical audio classification task, focusing specifically on Music Genre Classification (MGC). We develop and test three different Convolutional Neural Network (CNN) architectures, each containing various layer configurations, including Conv2d, Dropout, and Batch Normalization layers. These variations allow us to analyze the impact of each layer type on the overall model performance, providing insights into their individual contributions to effective genre classification. After that, we apply PROD fusion ensemble technique to combine multiple predicted probability vectors at inference phase to further improve the model performance. By conducting experiments on the best model performing an accuracy of $\mathbf{9 0 . 8 \%}$, which is a good accuracy and potential for real-life applications.