Chinese text retrieval without using a dictionary

Aitao Chen, Jianzhang He, Liangjie Xu, Fredric C. Gey, Jason Meggs · 1997

It is generafly believed that words, rather than characters, should be the smallest indexing unit for Chinese text retrieval systems, and that it is essential to have a comprehensive Chinese dictionary or lexicon for Chhmse text retrieval systems to do well.Chinese text has no delimiters to mark woni boundaries.As a result, any text retrieval systems that build word-based indexes need to segment text into words.We implemented several statistical and dictionary-hazed word segmentation methods to study the effect on retrieval effectiveness of different segmentation methods using the TREC-S Chinese test collection and topics.The results show that, for all three sets of queries, the simple bigram indexing and the purely statistical word segmentation perform better than the popular dictionary-based maximum matching method with a dictionary of 138,955 entries.

Read the paper · More papers on PaperTik