POS-tagging for non-english tweets: An automatic approach: (Study in Bahasa Indonesia)
Devi Munandar, Endang Suryawati, Dianadewi Riswantini, Achmad Fatchuttamam Abka, Rini Wijayanti, Andria Arisal · 2017
The studied approach to part-of-speech tagging for tweets in Bahasa Indonesia. Bahasa Indonesia, as well as many other non-English languages, are lacking in language processing resources. This is mainly due to the lack of usable local corpus. This paper describes our work on building tweet corpus in Bahasa Indonesia that have been tagged manually using tag set from previous work with the addition of new tags specifically for tweet data. This corpus then used to train neural network based tagger. The experiment results for training process obtained Skip-Gram model with 66.23% testing accuracy and 66.33% validation accuracy, f1 score for all model close 60% accuracy in POS tagger testing model and shows that the tagger can assign tags with the acceptable result using word vector as features.