Quantifying Train-Evaluation Overlap with Nearest Neighbors

Gauri Kambhatla, Nguyễn Thị Thanh Thủy, Eunsol Choi · 2023

Characterizing benchmark datasets is crucial to interpreting model performance.In this work, we study train-evaluation overlap as a measure of an individual dataset's adequacy to evaluate model generalization over a wide range of datasets.We quantify the overlap with a simple novel metric based on a nearest neighbors approach between the training and evaluation sets.We identify nearest training examples for each evaluation example by mapping instances with generic and task-specific embedding methods.Our study on eleven classification and extractive QA tasks reveals a wide range of trainevaluation overlap, and we show that the data collection method of the dataset and the difficulty of the task may play a role in the amount of overlap.Lastly, we use our nearest neighbor analysis to identify challenging or potentially mislabeled examples.Our analysis quantifies train-evaluation overlap, providing insights for constructing datasets to study generalization.

Read the paper · More papers on PaperTik