Sense-based clustering of Polish nouns in the extraction of semantic relatedness

Bartosz Broda, Maciej Piasecki, Stanisław Szpakowicz · Proceedings of the International Multiconference on Computer Science and Information Technology · 2008

The construction of a wordnet from scratch requires intelligent software support. An accurate measure of semantic relatedness can be used to extract groups of semantically close words from a corpus. Such groups help a lexicographer make decisions about synset membership and synset placement in the network. We have adapted to Polish the well-known algorithm of Clustering by Committee, and tested it on the largest Polish corpus available. The evaluation by way of a plWordNet-based synonymy test used Polish WordNet, a resource still under development. The results are consistent with a few benchmarks, but not encouraging enough yet to make a wordnet writer's support tool immediately useful.

Read the paper · More papers on PaperTik