A Human-Computuer Interaction Word Segmentation Method Adapting to Chinese Unknown Texts
Xiaohe Chen · Zhongwen xinxi xuebao · 2007
Word segmentation(WS)is a funamental task in Chinese information processing.To solve the difficulties of traditional methods in processing texts in restricted domains,a novel method is proposed.It requires no lexicon or training corpus and can adapt to various texts and different WS standards.It enables the user to take part in WS procedure and add language kownledge to the system.Using optimized suffix array algrithm,candidates as words are recursively extracted from the text,then judged and edited by the user.Thus,a lexicon of the text is gained and applied to segment the text.Experiments on 4 different texts show that without the user's judgement,F-score of the system reaches as much as 72%,and can be prompted by 12% with amount of work done by the user.With the increase in the workload of the user,the system is able to achieve better results.