Using Word Embeddings in Turkish Part of Speech Tagging

Şevket Can · International Journal of Machine Learning and Computing · 2021

The close relation between the stem (relatively the word meaning) and part of speech tag of the word turns part of speech tagging as an important preprocessing task in natural language processing and understanding problem.For example, if the Turkish word "gelecek" is labeled as noun, the word stem is to be "gelecek" meaning future.If it is labeled as verb, the stem is "gel" and in English it means, "come".In many languages including Turkish, part of speech tagging problem is generally solved by rule based approaches.In this paper, a setup where the neural network architecture SENNA together with word embeddings is employed.The combination of Wikipedia 2016 and METU corpora is utilized in training of word embeddings; PARDER is used in part of speech training and testing.The word embeddings that are obtained by different methods and different vector sizes are evaluated intrinsically considering analogic and semantic similarity distances; and assessed extrinsically based on the performance on part of speech tagging task.

Read the paper · More papers on PaperTik