All-words word sense disambiguation for Turkish

Onur Açıkgöz, Ali Tunca Gurkan, Burak Ertopçu, Ozan Topsakal, Berke Özenç, Ali Buğra Kanburoğlu, İlker Çam, Begüm Avar, Gökhan Ercan, Olcay Taner Yıldız · 2017 International Conference on Computer Science and Engineering (UBMK) · 2017

Identifying the sense of a word within a context is a challenging problem and has many applications in natural language processing. This assignment problem is called word sense disambiguation (WSD). Many papers in the literature focus on English language and data. Our dataset consists of 1400 sentences translated to Turkish from the Penn Treebank Corpus. This paper seeks to address and discuss 6 different feature extraction methods and its classification performances using C4.5, Random Forests, Rocchio, Naive Bayes, KNN, Linear and multilayer Perceptron. This paper calls into question how the described features perform on a morphologically rich language (Turkish) with several classifiers.

Read the paper · More papers on PaperTik