Methods for the Extraction of Hungarian Multi-word Lexemes

Balázs Kis, Gábor Pohl, Gábor Ugray, John A. Nerbonne · Repository of the Academy's Library (Library of the Hungarian Academy of Sciences) · 2004

This paper describes an experiment on extracting Hungarian multi-word lexemes from a corpus, using statistical methods.Corpus preparation-the addition of POS tags and stems-was done automatically.From the corpus, verb+noun+casemark patterns were extracted as collocation candidates.Evaluation shows that the statistical methods used by Villada Moirón (2004a) to identify Dutch V + PP collocations, can also be applied to the Hungarian data.Some collocation types (such as verbal arguments) require special extraction methods, as explained in the evaluation section.Finally, we suggest that the extraction process can be further improved by a blend of statistical techniques with rule-based and dictionary-based methods.

Read the paper · More papers on PaperTik