Leveraging Large Corpora Using Internet Search for Question Answering

Sean Gallagher, Wlodek Zadrozny · 2016

In this experiment, we measure the potential contribution of internet search to question answering. The task is to correctly answer Jeopardy! questions, and for doing so we use our existing question answering system, "Watsonsim", the architecture of which follows the original IBM Watson. We compare the answering precision and recall with and without internet search data from Microsoft's Bing. To accommodate the new sources we included several additional procedures, in particular merging candidate answers with similar meaning, and removing any candidate answers which are forbidden. The experiments show the potential of the internet search to improve question answering in terms of both the accuracy and recall: the top rank accuracy increased by around 14%, reaching 20% precision, compared to only 6% precision when using local search, and the recall improved by 20%, reaching 52%, up from 32% using only local search engines. Tweaks to better handle web results later raised top-rank accuracy an additional 21%, reaching 41% precision, and raised recall by another 19%, to reach 71% recall. The overall conclusion of our experiments was to show that in question answering, a network enabled system can compensate for a smaller corpus and provide reasonably good answers.

Read the paper · More papers on PaperTik