Evaluating Large Language Model Accuracy in Structured Academic Settings: Three Case Studies
Andres Fortino, Zoey Yang · 2024
As large language models (LLMs) are increasingly deployed in real-world applications, concerns exist regarding their accuracy and potential for hallucination. However, few studies evaluate performance systematically across academic contexts. This paper summarizes three case studies assessing LLM accuracy in classroom analytics exercises. In the first, students tested ChatGPT 3.5 responses to technical questions in graduate courses over 12 weeks. Over 95% factual concordance was found, with repetition being the primary inconsistency. The second study successfully replicated 167 statistical analyses from an Excel exercise book in ChatGPT 4 covering techniques like regression and inference. No discernible accuracy issues appeared. The third case evaluated Claude 2 and ChatGPT 4 on Harvard Business School cases and associated questions across four subjects, with student teams finding perfect alignment with faculty provided solutions. Collectively, the case studies reveal three key themes. First, constrained problem scopes limit deviations in accuracy. Second, performance relies on relevant domain exposure during pretraining. Third, high grading agreements position LLMs as credible teaching assistants, though personalization and originality remain inadequate. Contrary to wider benchmarks, LLMs demonstrate high fidelity on academic tasks, allaying concerns about hallucination. Speculatively, restricted interfaces could enhance reliability. Domain-specific finetuning also remains vital. The studies support LLMs' use for automated tutoring and analysis, complementing human educators. Further assessments on more advanced reasoning tasks are still necessary.