To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question Answering

Dheeru Dua, Emma Strubell, Sameer Kumar Singh, Pat Verga · 2023

Recent advances in open-domain question answering (ODQA) have demonstrated impressive accuracy on general-purpose domains like Wikipedia.While some work has been investigating how well ODQA models perform when tested for out-of-domain (OOD) generalization, these studies have been conducted only under conservative shifts in data distribution and typically focus on a single component (i.e., retriever or reader) rather than an end-to-end system.This work proposes a more realistic endto-end domain shift evaluation setting covering five diverse domains.We not only find that endto-end models fail to generalize but that high retrieval scores often still yield poor answer prediction accuracy.To address these failures, we investigate several interventions, in the form of data augmentations, for improving model adaption and use our evaluation set to elucidate the relationship between the efficacy of an intervention scheme and the particular type of dataset shifts we consider.We propose a generalizability test that estimates the type of shift in a target dataset without training a model in the target domain and that the type of shift is predictive of which data augmentation schemes will be effective for domain adaption.Overall, we find that these interventions increase end-to-end performance by up to ∼24 points.* *This work was done while authors were at Google.Average F1 over all target datasets Average F1 over target datasets with specific shifts

Read the paper · More papers on PaperTik