Generating Question Using Retrieval Augmented Generation with Large Language Model: Comparative Study
Nur Azmina Aisyah, Riyanarto Sarno, Abdullah Faqih Septiyanto, Agus Tri Haryono, Dwi Sunaryono · 2024
Rapid technological developments affect all aspects of life, especially in the field of education. One application of technology in the field of education is the automatic creation of questions. Automatic creation of questions is proposed by the author using the Retrieval Augmented Generation (RAG) method with Large Language Model (LLM). The aim is to help teachers in creating automatic exam questions by generating questions in the question bank or question database with the difficulty level of the questions. The author uses 3 test scenarios, namely ideal conditions, non-ideal conditions and slightly ideal conditions using 3 models, such as Llama 3, Cohere LLM and GPT 2. The author's goal in using 3 models is to compare models seen from the results of their accuracy. The results of the experiments conducted are that in ideal conditions, Llama-3-8B and GPT 2 get a very good accuracy score of 100%, while Cohere LLM gets an accuracy score of 80%. In non-ideal conditions, Cohere LLM gets the best score among other models. Cohere LLM gets an accuracy score of 80%, while Llama-3-8B and GPT 2 get an accuracy score of 67%. In slightly ideal conditions, Llama-3-8B gets the highest accuracy score of 100%. Cohere LLM and GPT 2 get a poor score of 67%. From these results, it can be concluded that Llama-3-8B is the best model compared to the other 2 models in generating questions. Then followed by GPT 2 and Cohere LLM