Towards speaker detection using lips movements for human-machine multiparty dialogue
Fasih Haider, Samer Al Moubayed · 2012
This paper explores the use of lips movements for the purpose of speaker and voice activity detection, a task that is essential in multi-modal multiparty human machine dialogue. The task aims at detecting who and when someone is speaking out of a set of persons. A multiparty dialogue consisting of 4 speakers is audiovisually recorded and then annotated for speaker and speech/silence segments. Lips movements are tracked using the real-time FaceAPI face tracking commercial software. The paper reports on results from 3 classification techniques, namely: neural networks, naïve Bayes classifiers, and Mahalanobis distance. In speech / silence detection, the experiments show promising results using lips movements with an optimal accuracy of 78.31%. The results also show that the neural network classifier has better results than other techniques in speaker dependent and hybrid method. However in speaker independent method the results show that the naïve Bayes classifier has the best result with accuracy of 64.56%.