Experiments on chinese text indexing : CLARIT TREC-5 Chinese track report
Tong Xiang, ChengXiang Zhai, Nataša Milić-Frayling, David Andreoff Evans · 1996
Introduction The focus of the CLARIT TM1 Chinese Track Experiments is on investigating the effectiveness of different automatic indexing methods for retrieval over Chinese texts. In particular, we explored indexing using linguistic units (words, compound words, and phrases), single Chinese characters, and overlapping character bigrams. In addition to fully automatic processing of queries, we ran experiments with manually constructed term vector queries supplemented by Boolean type constraints. The constraints were used for selecting documents for CLARIT automatic feedback or as a mean of refining the final set of retrieved documents [Mili'c-Frayling et al. 1997]. All the experiments were conducted using the CLARIT retrieval system [Evans & Lefferts 1995]. Since its current NLP component does not support the parsing of Chinese texts, we designed an appropriate parsing module and pre-processed the documents before submitting them for indexing and retrieval