English-French and English-Greek parallel corpus for the Environment and Labour Legislation domains

Pavel Pecina, Antonio Toral, Vassilis Papavassiliou, Prokopis Prokopidis, Victoria Arranz · Repositori digital de la UPF (Universitat Pompeu Fabra) · 2014

This report describes the deliverable D5.3 of the PANACEA project: parallel, sententially aligned texts, cleaned and prepared for training-building translation models. This deliverable is/nan outcome of the work packages WP4.1, WP4.2, and WP5.1. The resulting domain-specific parallel corpus includes a total of 4 million words for two language pairs: English–French and/nEnglish–Greek and two domains: Environment and Labour Legislation. The data for each language pair and domain is split into three sets: training data for building translation models, development data for development purposes, and test data for testing and evaluation.

Read the paper · More papers on PaperTik