Extraction of Polish noun senses from large corpora by means of clustering

Bartosz Broda, Maciej Piasecki, Stan Śzpakowicz · Control and Cybernetics · 2010

We investigate two methods of identifying noun sen- ses, based on clustering of lemmas and of documents. We have adapted to Polish the well-known algorithm of Clustering by Com- mittee, and tested it on very large Polish corpora. The evaluation by means of a WordNet-based synonymy test used Polish wordnet (plWordNet 1.0). Various clustering algorithms were analysed for the needs of extraction of document clusters as indicators of the senses of words which occur in them. The two approaches to word- sense identification have been compared, and conclusions drawn.

Read the paper · More papers on PaperTik