Approaches to visual speech processing based on the MPEG-4 Face Animation standard

Eric D. Petajan · 2002

Speech is communicated using both acoustic and visual modalities. Multi-modal automatic speech recognition (ASR) systems are not yet commercially available due to technical challenges in both the acquisition of visual speech features and the integration of the visual with the acoustic speech recognition processes. This paper explores the visual speech acquisition problem and describes how the MPEG-4 Face Animation standard efficiently represents visual speech information and encourages visual speech applications.

Read the paper · More papers on PaperTik