Handling multimodality and scarce resources in sign language machine translation

Christoph Schmidt · RWTH Publications (RWTH Aachen) · 2016

In the field of statistical machine translation, the translation of sign languages poses an interesting and challenging problem. Translating from a signed language into a spoken language is usually a two-step process. From the video of a person signing, an automatic sign language recognition system extracts the signs in some form. Since signed languages differ in grammar, vocabulary and expression from spoken languages even within the same country, the recognised signs have to be translated into a spoken language text.In sign languages, meaning is conveyed simultaneously not only via the two hands, but also by facial expressions, body posture, head movement, and eye gaze. Because of this complex and multimodal nature of sign languages, there is no common writing system, and the scientific question of an annotation scheme suitable for machine translation remains open. Another difficulty when applying statistical methods to sign language translation is the lack of a sufficient amount of training data. This data scarcity often leads to poor translation results. Moreover, the multimodal nature of sign language is not handled by current translation systems, which usually process sequences of words. In this thesis, we approach the above three problems: finding a suitable annotation scheme, dealing with small amounts of annotated data, and handling multimodality in the machine translation process. While we focus on the second step of translating the recognised signs into a spoken language text, we also aim at improving the overall process of recognition and translation by optimising the interface between the two systems. In the course of the EU-project SignSpeak, our research group implemented the whole pipeline of sign language recognition and translation, and we evaluated this pipeline with automatic measures and human evaluators.The annotation scheme of the RWTH-PHOENIX Weather 2014 Corpus, a sign language corpus which was created in the course of this thesis, contains not only information on the signs expressed by the hands, but also marks mouthings, locations in the signing space or simultaneous signing of two different signs with both hands. We analyse the importance of this additional information for machine translation and devise an improved way of including it in the process of translation.Since available sign language corpora are rather small when compared to spoken language corpora, the lack of sufficient training data often leads to a poor automatic alignment between the annotated signs and their translations in the spoken language. We improve the automatic alignment by applying a morphosyntactic and a semantic analysis and bridging the differences to find corresponding signs and phrases.To handle the multimodality of sign languages in statistical machine translation, we present two approaches, focusing on the hands and on mouthings, i.e. silent lip movements pronouncing certain words in the sentence, as two modalities. In the first approach, we automatically adapt the granularity of the annotation by distinguishing signs with the same hand movements but different mouthings based on an automatic extraction of lip movements. In the second approach, we use the mouthing directly in the decoding process, using both the information signed by the hands and the mouthing as an input to the decoder. By approaching the three issues of a suitable annotation, of data scarcity and of multimodality, we arrive at a translation system which can handle the multimodal sign language input and which is well beyond the performance of a standard translation system that only translates the manual component of a sign language utterance.

Read the paper · More papers on PaperTik