Building an Indonesian rule-based part-of-speech tagger
Rashel Fam, Andry Luthfi, Arawinda Dinakaramani, Hendra Manurung · 2014
This paper describes work on a part-of-speech tagger for the Indonesian language by employing a rule-based approach. The system tokenizes documents while also considering multi-word expressions and recognizes named entities. It then applies tags to every token, starting from closed-class words to open-class words and disambiguates the tags based on a set of manually defined rules. The system currently obtains an accuracy of 79% on a manually tagged corpus of roughly 250.000 tokens.