An Overview of Part-of-Speech Tagging Methods and Datasets for Malay Language

Chi Log Chua, Tong Ming Lim, Kwee Teck See · 2023

The purpose of this review is to summarise the knowledge about Malay Part-of-Speech (POS) training methods and datasets, and to identify its future research directions. A total of ten research papers related to Malay POS model training methods has been reviewed and five datasets were found from online resources. Two major issues were identified − first, of all the ten papers reviewed, it was found that dataset used to train Malay POS were not standardized. Second, limited dataset was found from online resources − only five datasets were available, and only three were annotated. This review highlights two directions worth future investigation: how to train a high-performance POS model using only a small amount of annotated data, and how to utilize existing high-performance POS models to reduce the burden of annotating data.

Read the paper · More papers on PaperTik