A novel Approach for Audio-based Video Analysis via MFCC Features
Ambreen Sabha, Arvind Selwal · Procedia Computer Science · 2024
The video analysis is a momentous task of computer vision where larger videos are analyzed and converted to comparatively shorter summaries with key-frames satisfying a typical criterion. However, in the recent times audio-based video analysis is an important field where, audio features are used to categorize video frames into various classes. In this work, we present an audio-based video analysis model named as ABVS. The proposed approach employs 40 MFCC audio features from input audios to learn a multiclass classifier. The ABVS model is trained on ANN, SVM, KNN, DT and a majority voting ensemble to categorize a given video based on audio at high precision and low error rate. Our proposed approach is trained and validated on a benchmark urbansound8k dataset with an accuracy of 90% and loss of 0.10 on ANN and SVM model whereas, KNN achieves an accuracy of 89%. The ABVS model can be deployed in real-time applications to extract key-frames from the video scenes to troubleshoot the malfunctioning of critical infrastructure in industrial sectors. Our proposed model may be deployed in audio forensics for criminal voice investigation.