Design considerations for developing a parts-of-speech tagset for Khasi

Medari Janai Tham · 2012

Several tagsets have been developed for Indian languages belonging to the Indo-Aryan and Dravidian families. This is because the major chunk of India's spoken language belongs to these categories. Khasi, on the other hand, belongs to the Austro-Asiatic family and is spoken primarily in the state of Meghalaya. To the best of my knowledge, language technology for Khasi is practically nonexistent and work on computational linguistic for the language is very scant. This proves to be a challenge when an attempt is made to provide access to technology using language when the basic tools needed are not available. There exists a common Part of Speech Tagset framework for Indian languages (IL-POSTS) covering the morphologically rich Indian languages under the Indo-Aryan and Dravidian families. However, in this paper the EAGLES guidelines are used for developing the Khasi tagset due to the natural infinity of the language to English. This is obvious from the script used, which is the Roman script and the word order is also primarily SVO.

Read the paper · More papers on PaperTik