Audiovisual speech recognition using multiscale nonlinear image decomposition
Iain A. Matthews, Jenny Bangham, Simón Cox · 2002
There has recently been increasing interest in the idea of enhancing speech recognition by the use of visual information derived from the face of the talker. This paper demonstrates the use of nonlinear image decomposition, in the form of a "sieve", applied to the task of visual speech recognition. Information derived from the mouth region is used in visual and audio-visual speech recognition of a database of the letters A-Z for four talkers. A scale histogram is generated directly from the gray-scale pixels of a window containing the talker's mouth on a per-frame basis. Results are presented for visual-only, audio-only and a simple audio-visual case.