Retrieving Collocations From Korean Text
Seonho Kim, Zooil Yang, Mansuk Song, Jungho Ahn · 1999
This paper describes a statistical methodology ibr automatically retrieving collocations from POS tagged Korean text using interrupted bigrams. The free order of Korean makes it hard to identify collocations. We devised four staffstics, 'frequency', 'randomness', 'condensation', and 'correlation' .to account for the more flexible word order properties of Korean collocations. Ve extracted meaningful bigrams using an evaluation function 'and extended the bigrams to n-gram collocatibns by generating equivalence sets, (-covers. We view a modeling problem for n-gram collocations as that for clustering of co- hesive words.