Marathi Word Sense Disambiguation through unsupervised K-Means Clustering

Rasika Ransing, Archana N. Gulati · Engineering Technology & Applied Science Research · 2025

Word Sense Disambiguation (WSD) is the most crucial Natural Language Processing task and refers to the process of determining the most suitable meaning of a word within its contextual usage. The case of the Marathi language is a bit complicated because it is considered a low-resource language, primarily due to the scarcity of annotated datasets. This study employs an unsupervised machine learning technique using k-means clustering for the disambiguation of Marathi words with more than one meanings without relying on manually labeled data. This disambiguation is accomplished with the help of the context these ambiguous words are used. Instead of implementing k-means clustering concurrently for all 12 words including 42 meanings, it is implemented separately for each word. The number of clusters for each word equals the number of meanings assigned to it. For each word, a Silhouette score is calculated to evaluate the quality of the obtained clustering. In the case of nouns, semantic boundaries were better defined, achieving higher Silhouette scores.

Read the paper · More papers on PaperTik