Using viseme recognition to improve a sign language translation system.
Christoph Schmidt, Oscar Koller, Hermann Ney, Thomas Hoyoux, Justus Piater · Open Repository and Bibliography (University of Liège) · 2013
Sign language-to-text translation systems are similar to spo-ken language translation systems in that they consist of a recognition phase and a translation phase. First, the video of a person signing is transformed into a transcription of the signs, which is then translated into the text of a spoken language. One distinctive feature of sign languages is their multi-modal nature, as they can express meaning simultane-ously via hand movements, body posture and facial expres-sions. In some sign languages, certain signs are accompanied by mouthings, i.e. the person silently pronounces the word while signing. In this work, we closely integrate a recog-nition and translation framework by adding a viseme recog-nizer (“lip reading system”) based on an active appearance model and by optimizing the recognition system to improve the translation output. The system outperforms the standard approach of separate recognition and translation. 1.