Visual speech recognition for isolated digits using discrete cosine transform and local binary pattern features
Abhilash Jain, G. N. Rathna · 2017
Visual Speech Recognition (VSR) deals with the task of extracting speech information from visual cues from a person's face while speaking. Accurate lip segmentation and modeling are essential in any VSR algorithm for good feature extraction. However, lip modeling is a complicated task and is not very robust in natural conditions. This paper describes a novel technique for limited vocabulary visual-only speech recognition that does not use lip modeling. For visual feature extraction, Discrete Cosine Transform (DCT) and Local Binary Pattern (LBP) have been tested. An Error-Correcting Output Codes (ECOC) multi-class model using Support Vector Machine (SVM) binary learners is used for recognition and classification of words.