A Preliminary Study on Taiwanese POS Taggers: Leveraging Chinese in the Absence of Taiwanese POS Annotation Datasets
Chao‐Yang Chang, Yan-Ming Lin, Chih-Chung Kuo, Yen-Chun Lai, Chao-Shih Huang, Yuan‐Fu Liao, Tsun-guan Thiann · 2024
Developing POS taggers for endangered and low-resource languages such as Taiwanese remains crucial due to the ex-tensive data demanded by modern LLMs like ChatGPT. This work focuses on leveraging the structural similarity between Chinese and Taiwanese through cross-language transfer learning. We adapt Chinese POS taggers to Taiwan-ese ones using two approaches: (1) modifying a CRF-based [1] Chinese POS tagger with a bilingual lexicon to map Taiwanese words to Chinese counterparts; (2) fine-tuning an LLM for Chinese POS tagging and using machine translation to transfer this capability to Taiwanese. Additionally, word and sentence-level contrastive learning is employed to create a shared semantic space between Chinese and Tai-wanese. The CRF-based Taiwanese POS tagging achieved an F1 score of 70.69% and the LLM approach achieved 69.72%. Our method aims to enable Taiwanese POS tagging without relying on extensive Taiwanese datasets, thus facili-tating the creation of a substantial Taiwanese POS tagging corpus with minimal manual effort.