Corpus Based Investigation on High Frequent Maximal Overlapping Ambiguity String in Chinese Word Segmentation
XU Yan-hua · Zhongwen xinxi xuebao · 2006
Overlapping ambiguity is still an open issue in Chinese word segmentation.This paper makes a deep investigation on Maximal Overlapping Ambiguity String(MOAS).First,we discuss the disadvantage of using FBMM to detect OAS.Then,by word omni-segmentation,we collect 14906 high frequent MOASs from People's Daily corpus which contains about 400M characters.For these MOASs,1354270 sample sentences are randomly selected and manually labeled.The results show that about 70% of MOASs with true ambiguity have a strong bias towards one segmentation,and consequently,a disambiguation strategy fon dealing with overlapping ambiguities is put forward.