A Comprehensive Survey of Datasets for Large Language Model Evaluation
Yuting Lü, Chao Sun, Yuchao Yan, Hegong Zhu, Dongdong Song, Qing Peng, Li Yu, Xiaozheng Wang, Jian Jiang, Xiaolong Ye · 2024
Natural Language Processing is an important branch of Artificial Intelligence. In the past few years, we have witnessed the remarkable advancement of large language models, however, how to evaluate them in a comprehensive way has become an urgent problem to be solved. Datasets can help evaluate and compare their performance and clarify their weaknesses. In order to guide the subsequent research work and promote the technological progress in the field, this paper collects 147 popular evaluation datasets, and proposes a new classification method to categorize them into six categories according to the evaluation capabilities. In addition, we organize several common evaluation metrics and usage scenarios. We compile the list of datasets and the main features (introduction, samples, metrics, links, etc.) into a document, which is consistently maintain available online at: https://github.com/lyt719/LLM-evaluation-datasets.