Rhythmic Resonance: A New Era of Musical Instrument Categorization with CNN-SVM
Ankita Suryavanshi, Shiva Mehta, Deepak Upadhyay, Manisha Aeri, Vishal Kumar Jain · 2024
This study investigates the efficiency of hybrid modelling that combines a Convolution Neural Network (CNN) and Support Vector Machine (SVM) to classify musical instrumentations in audio recordings. The categorization assignment is divided into five categories representing distinct families of musical instruments: strains, woodwinds, brass, percussions, and keyboard. The decision model's performance is rigorously measured using precision, recall, F1-score, and overall accuracy. The model value ranged between 91.26% and 95.02% concerning accuracy and 89.29% to 98.07% for recall across five class groups. We can conclude the F1 scores to be well-rounded in accuracy and recall, with values ranging from 91.20% to 95.42%, indicating excellence. Each class showed a broad variation of support with no example under or above the value 415 or 478, respectively, representing no support regions of 0.18 and 0.21. The model had the confidence of overall accuracy of up to 93.41%. Various types of averages were considered for a complete discussion of the model's performance. The macro average accuracy was 93.38%, and recall and F1 scores were 93.41% and 93.52%, respectively. The weighted average accuracy was 93.43%, the recall was 93.41%, and the F1-score was 93.39%, although considering class imbalance. The microaverage, representing all classes' net contribution to accuracy, recall, and the F1-score, remained at 93.41%. The experimental results demonstrate that the CNN-SVM model is a strong and capable instrument identification model in complex acoustics.