A LARGE SPANISH-CATALAN PARALLEL CORPUS RELEASE FOR MACHINE TRANSLATION

Marta R. Costa‐jussà, José A. R. Fonollosa, José Bernardo Mariño Acebal, Marc Poch, Mireia Farrús · Repositori digital de la UPF (Universitat Pompeu Fabra) · 2014

We present a large Spanish-Catalan parallel corpus extracted from ten years of the paper edition of a bilingual Catalan newspaper. The produced corpus of 7:5M parallel sentences (around 180M words per language) is useful for many natural language applications. We report excellent results when building a statistical machine translation system trained on this parallel corpus. The Spanish-Catalan corpus is partially available via ELDA (Evaluations and Language Resources Distribution Agency) in catalog number ELRA-W0053.

Read the paper · More papers on PaperTik