The OPUS corpus : parallel and free

Jörg Tiedemann, Lars Nygard · 2004

The OPUS corpus is a growing collection of translated documents collected from the internet. The current version contains about 30 million words in 60 languages. The entire corpus is sentence aligned and it also contains linguistic markup for certain languages. 1.

Read the paper · More papers on PaperTik