The Hungarian Gigaword Corpus

Csaba Oravecz, Tamás Váradi, Bálint Sass · 2014

The paper reports on the development of the Hungarian Gigaword Corpus, an extended new edition of the Hungarian National Corpus, with upgraded and redesigned linguistic annotation and an increased size of 1.5 billion tokens.Issues concerning the standard steps of corpus collection and preparation are discussed with special emphasis on linguistic analysis and annotation due to Hungarian having some challenging characteristics with respect to computational processing.

Read the paper · More papers on PaperTik