Human-Like e-Learning Mediation Agents
Chukwuka Victor Obionwu, Diptesh Mukherjee, Andreas Nürnberger, Aarathi Vijayachandran Bhagavathi, Aishwarya Harneer Suresh, Eathorne Choongo, Bhavya Baburaj Chovatta Valappil, Amit Kumar, Gunter Saake · 2025
The field of education has undergone a significant transformation in recent years, largely due to the ongoing development in the areas of natural language processing and generative artificial intelligence (AI). This has given rise to digital platforms that are capable of facilitating the mediation of students learning interactions in ways that were previously unattainable. These advancements have not only enhanced the efficiency of educational processes but have also opened up new possibilities for personalized and adaptive learning experiences. At the forefront of this transformation are conversational agents, knowledge graphs, and large language models, each representing a distinct facet of the ed-tech revolution. These tools and technologies have not only allowed for new ways of presenting information and assessing student progress, but have also revolutionized the way educators can provide individualized support and feedback to students. In our particular scenario, we have sought to integrate human-like task generation, evaluation, and instructional feedback into our online learning platform to enhance the student experience and promote a deeper understanding of the material. While previous efforts have mostly been retrieval centric systems, the shift to a retrieval-augmented strategy has resulted in the provision of feedback that rivals the baseline performance of human instructors in both task generation, evaluation, and feedback augmentation. In this work, we detail the current strategy and performance details of our retrieval-augmented chat system and the task generation and evaluation system. Our evaluation methodology uses a range of automatic evaluation metrics, including BLEU, METEOR, ROUGE, and BERT similarity scores. As will be shown in subsequent sections, T5-Small demonstrates noteworthy performance despite having comparatively fewer parameters compared to GPT and T5-Large, which we used as baselines. Furthermore, we delved into the correlation between these automated metrics and human judgment by comparing the model&s;s scores against human-assigned ratings. The evaluation results reveal that both the BERT similarity score and the Meteor score exhibit a notably higher degree of similarity with the assessments provided by human evaluators.