Evaluating the use of word embeddings for part-of-speech tagging in Bahasa Indonesia
Achmad Fatchuttamam Abka · 2016
This paper studies the use of word embeddings for POS tagging in Bahasa Indonesia. The experiments are conducted with an architecture based on neural network model, that is a simple feed forward neural network with one hidden layer. The word embeddings (i.e., CBOW, skip-gram, and GloVe) are trained on unlabelled text corpus created from Wikipedia Bahasa Indonesia. The results show that word embeddings can be used for POS tagging in Bahasas Indonesia with good performance. The F1 score of all word embeddings types are roughly 80% and the accuracies are higher than 93% on a manually tagged corpus of about 250.000 tokens (12,775 unique tokens).