Role of collocations and case-markers in word sense disambiguation: a clustering-based approach
Baskaran Sankaran, V. Vaidebi · 2005
In this paper we present a new unsupervised approach for Word Sense Disambiguation (WSD) based on clustering. We also explore the applicability of case markers in WSD for inflectional languages such as Tamil. All the occurrences of an ambiguous word are grouped into as many unique clusters as the number of senses of the ambiguous word in such a way that each cluster contains occurrences in a single sense. After clustering the occurrences, highly reliable collocation words indicating strongly one particular sense of the ambiguous word are automatically selected from each cluster. Then for each set of collocations obtained, senses are assigned manually. This approach when compared to Decision tree approach (Yarowsky 1995) requires less manual labour. In our approach, collocations are induced automatically, and manual efforts are needed only in assigning the senses to different set of collocations, which is much easier task compared to Yarowsky’s method. The hypothesis for case-marker based disambiguation is that "Each sense of the ambiguous word takes different case markers in its neighbourhood". Our preliminary analysis proves the worthiness of this property and this method combined with the clustering based approach performs well in disambiguating the senses. Though the tested accuracy is around 84 % now, we are hopeful of approaching the accuracy levels of (Yarowsky 1995).