Automatic Audio Classification and Speaker Identification for Video Content Analysis

Shu-Chang Liu, Jing Cun Bi, Zhiqiang Jia, Rui Chen, Jie Chen, Min-Min Zhou · 2007

Recently, more literatures proposed to apply audio content analysis techniques in content-based video parsing. This paper presents our works on audio classification and speaker identification techniques for video content analysis. Firstly, soundtrack extracted from video stream is partitioned into homogeneous segments using rule and Support Vector Machine(SVM) based classifier. Secondly, fixed-length speech clips randomly selected from speech segments are clustered into several clusters based on spectral clustering techniques. The clustered speech feature datasets initialize and train Gaussian Mixture Model(GMM) for each speaker. Finally, the trained GMMs accomplish speaker identification. Experimental results confirm the validity of the proposed scheme.

Read the paper · More papers on PaperTik