ORCHID: Thai Part-Of-Speech Tagged Corpus
Virach Sornlertlamvanich, Thatsanee Charoenporn, Hitoshi Isahara · 2009
This paper presents a procedure in building a Thai part-of-speech (POS) tagged corpus named ORCHID [1]. It is a collaboration project between Communications Research Laboratory (CRL) of Japan and National Electronics and Computer Technology Center (NECTEC) of Thailand. We proposed a new tagset based on the previous research on Thai parts-of-speech for using in a multi-lingual machine translation project. We marked the corpus in three levels:paragraph, sentence and word. The corpus keeps text information in text information line and numbering line which are necessary in retrieving process. Since there are no explicit word/sentence boundary, punctuation and inflection in Thai text, we have to separate a paragraph into sentences before tagging the POS. We applied a probabilistic trigram model for simultaneously word segmenting and POS tagging. Rule for syllable construction is additionally used to reduce the number of candidates for computing the probability. The problems in POS assignment are formalized to reduce the ambiguity occurring in case of the similar POSs.