Selecting the best feature set for Thai word sense disambiguation using support vector machines

Chutchada Nusai, Yoshimi Suzuki, Haruaki Yamazaki · Conference on Artificial Intelligence for Applications · 2007

This paper proposes a method of selecting the best feature set for Thai sense disambiguation by using Support Vector Machines (SVM) algorithm. This research focuses on Thai verb sense disambiguation. Many approaches have been employed to resolve the sense ambiguity with a reasonable degree of accuracy. Our research focuses on the corpus-based approach that employs a supervised machine learning method for disambiguation. The machine learning method has the ability of selecting the suitable feature. In order to find the best feature set for resolving Thai sense ambiguity, our method uses characteristics of the words co-occur with the ambiguous in sentences extracted from Thai corpus for determining sense of the ambiguous word. The ambiguous words are evaluated with 30 feature sets under part of speech (POS) and semantic concept (SM) features. The result shows that word & SM feature set gives the best result as the best feature set of sense indicator and the accuracy rate is approximately 90-96%.

Read the paper · More papers on PaperTik