NTCIR-6 Monolingual Chinese and English-Chinese Cross-Lingual Question Answering Experiments using PIRCS
Kui Lam Kwok, Peter Deng, Norbert Dinstl · 2007
We continue to employ a minimal approach for our Chinese QA work that requires only a COTS entity extraction software and other home-built tools. In monolingual Chinese QA, questions are classified based on cue-word and meta-keyword usage patterns. Retrieval is done using sentence units, and indexing is based on bigrams and characters. Entities extracted from retrieved sentences form a pool of answer candidates which are ranked using five evidence factors. Our best monolingual result shows that when only Top1 answers are considered, 63 questions out of 150 are answered correctly with sentence support, giving an accuracy and MRR of 0.42. When unsupported answers are included, these values improve to 0.4467. English-Chinese CLQA starts with English question classification also based on an approach similar to Chinese. Three paths of translation render the question into Chinese strings. Otherwise procedures of retrieval and answer ranking remain the same as monolingual but with different parameter values. Our best run returns corresponding Top1 values as:.2533 and.28 (unsupported). These are about 60 % of monolingual effectiveness within our system. Effectiveness with Top2-5 answers as well as the influence of different evidence factors are also reported.