Linguistic characterization of answer passages for fact-seeking question answering

Bahadorreza Ofoghi · Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing · 2022

The performance of the passage ranking process for Question Answering (QA) systems has seen significant improvements, especially with recent advances in deep learning and distributed textual representation mechanisms. As the effectiveness of rankers and comprehension systems has improved, our understanding of what constitutes specific answer-containing passages has been shadowed under the complex inner mechanics of deep learning constructs. In this paper, rather than making attempts at developing a yet new deep learning structure for answer passage ranking, we step back on several of the core linguistic characteristics of questions and relevant passages to gain a more comprehensive understanding of answer-containing text snippets in relation to fact-seeking natural language questions. We have experimented with two large benchmark data sets, QUASAR-T and TriviaQA, and extracted several linguistic features from the questions and passages. Our statistical and Bayesian analyses show that answer-containing and answer-free passages have significantly different distributions of linguistic features and that a Naïve Bayes model can effectively capture, model, and explain feature relationships to distinguish between answer-containing and answer-free passages.

Read the paper · More papers on PaperTik