Designing relevant features for visual speech recognition

Eric Benhaim, Hichem Sahbi, Guillaume Vitte · 2013

Automatic speech analysis is currently evolving towards hybrid systems that combine both visual and acoustic information. This is due to limitations of existing acoustic-based approaches and the need for robust speech recognition systems working under extremely challenging conditions including noisy environments. We introduce in this paper a novel visual speech recognition approach, based on string kernels and support vector machines. The main contributions of this work include (i) the design of a similarity function, based on string kernels, that models the dynamics as well as the appearance of visual features in talking faces and (ii) a kernel combination procedure based on multiple kernel learning, that makes visual feature selection effective and also more tractable. Experiments conducted, on a standard digit database, show that the proposed algorithm outperforms current state-of-the-art methods.

Read the paper · More papers on PaperTik