Evaluating Chinese Large Language Models on Discipline Knowledge Acquisition via Memorization and Robustness Assessment

Chuang Liu, Renren Jin, Mark J. Steedman, Deyi Xiong · 2024

Chinese large language models (LLMs) demonstrate impressive performance on NLP tasks, particularly on discipline knowledge benchmarks, where certain Chinese LLMs are very competitive to GPT-4.Previous research has viewed these advancements as potential outcomes of data contamination or leakage, prompting efforts to create new detection methods and address evaluation issues in LLM benchmarks.However, there has been a lack of comprehensive assessment of the evolution of Chinese LLMs.To bridge this gap, this paper offers a thorough investigation of Chinese LLMs on discipline knowledge evaluation, delving into the advancements of various LLMs, including a group of related models and others.Specifically, we have conducted six assessments ranging from knowledge memorization to comprehension for robustness, encompassing tasks like predicting incomplete questions and options, identifying behaviors by the contaminational fine-tuning, and answering rephrased questions.Experimental findings indicate a positive correlation between the release time of LLMs and their memorization capabilities, but they struggle with variations in original question-options pairs.Additionally, our findings suggest that question descriptions have a more significant impact on the performance of LLMs.

Read the paper · More papers on PaperTik