Multilingual word segmentation and part-of-speech tagging : a machine learning approach incorporating diverse features
Tetsuji Nakagawa · Institutional Repositories DataBase (IRDB) · 2006
The aim of this dissertation is to study statistical methods for multilingual word segmentation and POS tagging with high accuracy.Word segmentation and part-of-speech (POS) tagging are fundamental language analysis tasks in natural language processing, and used in many applications.Existence of unknown words is a large problem in these tasks and they need to be properly handled.We attempt to develop suitable methods for word segmentation and POS tagging which can utilize informative features effectively.Firstly, we study a method for unknown word guessing and part-of-speech tagging using support vector machines (SVMs), which can handle a number of features effectively.We apply the method to English unknown word guessing and part-of-speech tagging.Secondly, we propose a method for POS guessing of unknown words using global information as well as local information.Global features often give useful information for POS guessing, and the method takes into consideration interactions between the POS tags of all the unknown words in a document by using Gibbs sampling.We apply the method to Chinese, Japanese and English unknown word guessing.Thirdly, we propose a word segmentation method which combines the existing word-based method and character-based method, in order to compensate for the