Exploring the Potential of Using AI Language Models in Democratising Global Language Test Preparation
Amalia Novita Sari · International Journal of TESOL & Education · 2024
This paper delves into the potential of AI language models for democratising global language test preparation, focusing on the accuracy and consistency of assessment in the context of writing essays for IELTS. This quantitative study compares the assessment scores generated by a Human Examiner (HE) and four AI Language Models: ChatGPT, Google Bard, Writing9.com, and Upscore.ai. Evaluation uses Mean Absolute Errors (MEA) and Bland Altman analysis. The findings reveal varying levels of accuracy, with Upscore.ai showcasing the lowest MEA of 0.5, followed by Google Bard at 0.85, ChatGPT at 0.9, and Writing9.com at 1.9. Bland Altman Plots visually represent the agreements between each alternative evaluation system and the Human Examiner, shedding light on their alignment. These results hold significant implications for assisting IELTS test takers in their preparation and advancing the democratisation of IELTS and global language assessment by harnessing AI technology to provide more accessible evaluation methods. AI evaluation systems can support teaching and learning by providing automated feedback when human assistance is unavailable, helping students practice independently. However, the findings show that AI's accuracy is not absolute and varies between models, meaning human involvement remains crucial for comprehensive evaluation.