Student use of accurate and inaccurate chatbot content: an empirical study
John Lee · 2025
With the integration of generative artificial intelligence (AI) into education, high-quality materials generated by Large Language Models (LLMs) have been shown to bring pedagogical benefits. Although it is well known that these models can hallucinate, there has been relatively less empirical research on the impact of misinformation on students’ learning outcome. This paper investigates whether university students can engage critically with LLMs and, specifically, the extent to which they can both benefit from accurate LLM content and recognize inaccurate content. In our study, 144 students answered short questions that required them to compare or distinguish between two concepts, scores or corpus queries. The correct answer may be one of the two options, or “both”. The answers generated by a chatbot were also shown to the treatment group, but not to the control group. In questions where the chatbot was correct, the treatment group outperformed the control group. In questions where the chatbot was incorrect, student performance varied according to the content of the chatbot answer. When the answer should be “both” but the chatbot accepted only one of the two options, the treatment group was more likely than the control group to recognize the validity of both options. However, when the chatbot also argued that the other option was incorrect, the treatment group was more prone to agree with the chatbot. Educators may find these results helpful in preparing students for the use of chatbot in their studies.