Okapi Chinese text retrieval experiments at TREC-6
Jimmy Xiangji Huang, Stephen E. Robertson · 1997
Introduction The focus of the Okapi TREC--6 Chinese experiments is on investigating the effectiveness of different automatic indexing methods and phrase weighting for retrieval based on probabilistic models over Chinese text. We compare different probabilistic weighting methods based on a range of word and single character approaches. There are two indexing methods used in our experiments. One indexing method is to use linguistic units (words, compound words and phrases) in texts as indexing terms to represent the texts. We refer to this method as the word approach. For this approach, text segmentation, which divides text into linguistic units, is regarded not only as a necessary precursor but also as a bottleneck of this kind of system [1]. The other method for indexing texts is based on single Chinese characters, in which texts are indexed by the characters appearing in the texts [2]. By using single character approaches, a search could be conducted for any multi-character w