EFFECTIVE PASSAGE RETRIEVAL IN QUESTION ANSWERING SYSTEMS

Surya Ganesh · 2010

Information Retrieval systems like Web search engines are often used by people to find information of their interest. Despite the success of these systems, they are not sophisticated enough to provide precise information for users requests. To address this problem, complex Information Access systems called Question Answering systems have been developed. They aim at finding exact answers to natural language questions in a large collection of documents (such as World Wide Web). One of the most obvious limitations in the performance of many Question Answering systems is their inability to find text passages where candidate answers can be found. Earlier research on the poor performance of passage retrieval highlighted the terminological gap problem i.e., passages holding the answer to a question have semantic alterations of original terms in the question. In this thesis, we proposed two different techniques to reduce this problem. Query expansion is a widely used technique in Information Retrieval to reduce the terminological gap problem. First, we present a novel passage retrieval methodology which expands the query inherently. This methodology leverages Statistical Machine Translation model for Information Retrieval to retrieve a ranked set of passages given a question. The retrieval within this model includes two steps: estimation and ranking. In the estimation step, multiple translation models are constructed using a statistical alignment model. We perceive each such translation model as an answer type profile. During ranking, based on the answer type of the question its corresponding answer type profile is used to retrieve relevant passages. Our experimental analysis on the performance of this retrieval methodology showed significant improvements over different standard retrieval models including TFIDF, Okapi BM25, Indri and KL-divergence. We found that simple statistical alignment models like IBM model 1 are more suited for passage retrieval. We also showed that our methodology addresses the problems of synonymy and polysemy. Previous studies on explicit query expansion methods like pseudo relevance feedback, and methods based on external knowledge sources likeWordNet, Wikipedia or Web have shown to improve the performance of Information Retrieval systems. We proposed a novel query expansion method using Wikipedia. Our methodology uses text content, category structure, and link structure of Wikipedia to generate a set of terms semantically related to the question. As Boolean model allows a finegrained control over query expansion, these semantically related terms are added to the original query to form an expanded Boolean query. Our experimental analysis on the performance of these expanded queries on Lucene, an open source retrieval system, showed significant improvements over seed queries. We also analyzed the performance of expanded queries based on different scoring methods utilized in selecting query expansion terms and for different query expansion lengths. Adding to the above contributions which focused on reducing the terminological gap between query and passage, we explored the necessity of passage priors in ranking passages. Document Retrieval assumes that a document is independent of its relevance, and non-relevance. The same assumption is being carried forward in passage retrieval systems in the context of Question Answering. We relax this assumption and explore the necessity of passage priors being relevant and nonrelevant in ranking passages given a query. We describe a mutual information measure for estimating these priors and a simple method for identifying relevant and non-relevant text to a question using the Web and AQUAINT corpus as information sources. Our experimental analysis of using passage priors as a re-ranking step on top of language models including Indri and KL-divergence showed that passage priors are necessary in ranking passages.

Read the paper · More papers on PaperTik