SUMAT: Data Collection and Parallel Corpus Compilation for Machine Translation of Subtitles
Volha Petukhova, Rodrigo Agerri, Mark Fishel, Sergio Penkale, Arantza del Pozo, Mirjam Sepesy Maučec, Andy Way, Panayota Georgakopoulou, Martin Volk · 2012
This paper describes the data collection and parallel corpus compilation activities carried out in the FP7 EU-funded SUMAT project.This project aims to develop an online subtitle translation service for nine European languages combined into 14 different language pairs.This data provides bilingual and monolingual training data for statistical machine translation engines which will semi-automate the subtitle translation processes of subtitling companies on a large scale.