Towards Assessing the Credibility of Chatbot Responses for Technical Assessments in Higher Education

Ritwik Murali, Dhanya M. Dhanalakshmy, Veeramanohar Avudaiappan, Gayathri Sivakumar · 2024

The recent challenge in higher education is to convey the importance of understanding concepts over rote learning. This challenge has increased in complexity with the arrival of large language model (LLM) based chatbots. Students are increasingly looking to such AI based chatbots as “sources of wisdom” instead of utilizing the same as learning aids. Despite disclaimers by the LLM creators, many students turn to the chatbot for answers to almost all learning assignments. This research work explores the level to which the LLM responses can be utilized for student learning in technical education. By understanding the contradictions between student answers and the responses generated by the LLMs, this work explores the limitations of the LLM based environments towards providing acceptable answers for assessments - specifically within the computer science engineering domain. While numerous studies have concentrated on ChatGPT, it is essential to consider the diverse range of alternative chat-bots accessible online that students may also utilize. Therefore, this work considers 5 popular AI-based chatbots for the study. With the “prompt” being the prime factor that impacts the response from chat-bots, the responses of the chatbots were collected using 2 different prompting techniques. The chatbot responses were evaluated against actual student responses by multiple reviewers to gauge its effectiveness as appropriate student answers. Both students and all chatbots were given questions aligned with the Blooms taxonomy levels (BTL) 1 to 4 in three different subjects. Each of the courses included a diverse range of questions including text-based questions, mathematical problems, and programming questions. The results show that the chatbot responses were acceptable for low BT level questions but failed to answer convincingly when asked for an algorithm. Overall, the chatbot performance (across the tested LLMs) was below average when the question set covered the BTL range 1–4. However, since the answers up to BTL2 were acceptable, LLM based chatbot answers were able to barely pass 1–2 of the 3 subjects (with the best performers scoring near the pass mark). Based on these results, it is possible to conclude that LLM based chatbots cannot be depended on for higher order learning but can be used to aid students who are struggling to pass basic courses.

Read the paper · More papers on PaperTik