Collocation based word sense disambiguation using clustering for Tamil

Baskaran Sankaran, V. Vaidehi · IJDL. International journal of Dravidian linguistics · 2004

In this paper we present an unsupervised approach for Word Sense Disambiguation using clustering technique. This approach relies on the automatically extracted collocations for tagging the ambiguous words with the appropriate senses. The advantage of this approach is that the dependence of word sense disambiguisation algorithms, on either a sense-tagged corpus or a knowledge-source, such as dictionary, thesaurus etc. is removed. The algorithm is first trained on a training corpus to extract relevant collocations for each ambiguous words. We use the clustering technique to group the occurrences of an ambiguous word in the training corpus, where each group corresponds to a particular sense of the ambiguous word. Prominent collocations are then identified for each group. These collocations are then tagged with the appropriate sense manually and this information is used for disambiguating unseen occurrences of the ambiguous word. This method reduces the human effort in terms of sense tagging compared to the ealier approaches and at the same time gives reasonable levels of accuracy.

Read the paper · More papers on PaperTik