The Use of Sentence Similarity as a Semantic Relevance Metric for Question Answering.
Marco De Boni, Suresh Manandhar · 2003
An algorithm for calculating semantic similarity between sentences using a variety of linguistic information is presented and applied to the problem of Question Answering. This semantic similarity measure is used in order to determine the semantic relevance of an answer in respect to a question. The algorithm is evaluated against the TREC Question Answering test-bed and is shown to be useful in determining possible answers to a question. Not all linguistic information is shown to be useful, however, and an in-depth analysis shows that certain sentence features are more important than others in determining relevance. Semantic relevance for Question Answering Question Answering Systems aim to determine an answer to a question by searching for a response in a collection of documents (see Voorhees 2002 for an overview of current systems). In order to achieve this (see for example Harabagiu et al. 2002), systems narrow down the search by using information retrieval techniques to select a subset of documents, or paragraphs within documents, containing keywords from the question and a concept which corresponds to the correct question type (e.g. a question starting with the word “Who?” would require an answer containing a person). The exact answer sentence is then sought by either attempting to unify the answer semantically with the question, through some kind of logical transformation (e.g. Moldovan and Rus, 2001) or by some form of pattern matching (e.g. Soubbotin 2002; Harabagiu et al. 2002). Semantic relevance (Berg 1991) is the idea that an answer can be deemed to be relevant in respect to a question in virtue of the meanings of the words in the answer and question sentences. So, for example, given the question Q: What are we having for dinner? And the possible answers