WMT16 APE Shared Task Data

Marco Turchi, Rajen Chatterjee, Matteo Negri · Americanae (AECID Library) · 2016

Training, development and text data (the same used for the Sentence-level Quality Estimation task) consist in English-German triplets (source, target and post-edit) belonging to the IT domain and already tokenized. Training and development respectively contain 12,000 and 1,000 triplets, while the test set 2,000 instances. All data is provided by the EU project QT21 (http://www.qt21.eu/).

Read the paper · More papers on PaperTik