A guide to the dataset explosion in QA, NLI, and commonsense reasoning
Anna Rogers, Anna Rumshisky · 2020
Question answering, natural language inference and commonsense reasoning are increasingly popular as general NLP system benchmarks, driving both modeling and dataset work.Only for question answering we already have over 100 datasets, with over 40 published after 2018.However, most new datasets get "solved" soon after publication, and this is largely due not to the verbal reasoning capabilities of our models, but to annotation artifacts and shallow cues in the data that they can exploit.The tutorial will be of the cutting-edge type.The tutorial slides are available online at https://www. annargrs.github.io/dataset-explosion.Prerequisites.We assume basic familiarity with the standard machine learning evaluation workflow and the three tasks that we are covering (question answering, commonsense reasoning, natural language inference).We also assume some familiarity with the methodology of crowdsourcing NLP datasets.