Data-intensive Question Answering

Eric Brill, Jimmy Lin, Michele Banko, Susan Dumais, Andrew Y. Ng · 2001

Data-driven methods have proven to be powerful techniques for natural language processing. It is still unclear to what extent this success can be attributed to specific techniques, versus simply the data itself. For example, in [Banko and Brill 2001] it was demonstrated that for confusion set disambiguation, a prototypical disambiguation-instring -context problem, the amount of data used far dominates the learning method employed in improving labeling accuracy. The more training data that is used, the greater the chance that a new sample being processed can be trivially related to samples appearing in the training data, thereby lessening the need for any complex reasoning that may be beneficial in cases of sparse training data. The idea of allowing the data, instead of the methods, do most of the work is what motivated our particular approach to the TREC Question Answering task. One of the biggest challenges in TREC-style QA is overcoming the surface string mismatch betw

Read the paper · More papers on PaperTik