A novel Approach for Audio-based Video Analysis via MFCC Features

Ambreen Sabha, Arvind Selwal · Procedia Computer Science · 2024

The video analysis is a momentous task of computer vision where larger videos are analyzed and converted to comparatively shorter summaries with key-frames satisfying a typical criterion. However, in the recent times audio-based video analysis is an important field where, audio features are used to categorize video frames into various classes. In this work, we present an audio-based video analysis model named as ABVS. The proposed approach employs 40 MFCC audio features from input audios to learn a multiclass classifier. The ABVS model is trained on ANN, SVM, KNN, DT and a majority voting ensemble to categorize a given video based on audio at high precision and low error rate. Our proposed approach is trained and validated on a benchmark urbansound8k dataset with an accuracy of 90% and loss of 0.10 on ANN and SVM model whereas, KNN achieves an accuracy of 89%. The ABVS model can be deployed in real-time applications to extract key-frames from the video scenes to troubleshoot the malfunctioning of critical infrastructure in industrial sectors. Our proposed model may be deployed in audio forensics for criminal voice investigation.

Read the paper · More papers on PaperTik