MoQA: Benchmarking Multi-Type Open-Domain Question Answering

Howard Yen, Tianyu Gao, Jinhyuk Lee, Danqi Chen · 2023

Previous research on open-domain question answering (QA) focuses mainly on shortanswered questions.However, informationseeking QA often requires various formats of answers depending on the nature of the questions, e.g., why/how questions typically require a long answer.In this paper, we present MOQA 1 , a benchmark for opendomain QA that requires building one system that can provide short, medium, long, and yes/no answers to different questions accordingly.MOQA builds upon Natural Questions (Kwiatkowski et al., 2019) with multiple types of questions and additional crowdsourcing efforts to ensure high data quality.We adapt state-of-the-art models, and reveal unique findings in multi-type open-domain QA: (1) For retriever-reader models, training one retriever on all types achieves the overall best performance, but it is challenging to train one reader model to output answers of different formats, or to train a question classifier to distinguish between types; (2) An end-to-end closed-book QA model trained on multiple types struggles with the task across the board; (3) State-of-theart large language models such as the largest GPT-3 models (Brown et al., 2020;Ouyang et al., 2022) also lag behind open-book QA models.Our benchmark and analysis call for more effort to build versatile open-domain QA models in the future.2

Read the paper · More papers on PaperTik