Coling 2008: Proceedings of the 2nd workshop on Information Retrieval for Question Answering
Mark Greenwood · 2008
Open domain question answering (QA) has become a very active research area over the past decade, due in large measure to the stimulus of the TREC Question Answering track (now a track within the recently formed Text Analysis Conference, TAC). This track addresses the task of finding answers to natural language questions (e.g. How tall is the Tower?, Who is Aaron Copland?, What effect does second-hand smoke have on non-smokers?) from large text collections. This task stands in contrast to the more conventional information retrieval (IR) task of finding documents relevant to a query, where the query may be simply a collection of keywords (e.g. Eiffel Tower, American composer, born Brooklyn NY 1900, ...). Finding answers requires processing texts at a level of detail that cannot be carried out at retrieval time for very large text collections. This limitation has led many researchers to rely on, broadly, a two stage approach to the QA task. In stage one a subset of question-relevant texts are selected from the whole collection. In stage two this subset is subjected to detailed processing for answer extraction. Clearly performance at stage two is bounded by performance at stage one, and previous work has shown that, despite the sophistication of standard IR ranking algorithms, they are not well suited to the stage one task of retrieving relevant documents given short natural language questions. It is likely that improvements in this area will come from linguistic insights into why QA focused IR is different from the traditional IR model. With the continued expansion of QA research into more complex question types and with the speed with which answers are returned becoming an issue, the importance of having good, QA-focused IR techniques is likely to increase. To date this topic has received limited explicit attention despite its obvious importance. This 2nd IR4QA workshop aims to address this situation by continuing to attract the attention of researchers to the specific IR challenges raised by QA.