The MULTEXT East corpus

Tomaž Erjavec, Nancy Ide · 1998

The EU MULTEXT-East project has produced harmonised language resources for Bulgarian, Czech, Estonian, Hungarian, Roma-nian, and Slovene. In this paper we introduce the MULTEXT-East multilingual corpus, which comprises marked-up texts in the six languages totaling approximately 2 million words and a small speech corpus. The corpus is encoded in SGML, in the TEI-like Corpus Encoding Specification and is divided into a parallel and a comparable (fiction, news) part. The parallel corpus consists of the novel "1984 " by George Orwell in the English original and translations. The translations are sentence aligned with the original and tagged for word-level linguistic information, i.e. for morphosyntactic descriptions and lemmas. Detailed information on the corpus is available on the WWW and the corpus itself has been released for research purposes on a CD-ROM in the scope of the TELRI concerted action. 1. Overview While standardised, large-scale language resources ex-ist or are under development for most western languages there have, so far, been few comparable efforts for Central and Eastern European (CEE) languages. The MULTEXT-East (Multilingual Text Tools and Corpora for Eastern and Central European Languages) project (Erjavec et al., 1996)

Read the paper · More papers on PaperTik