Generation of Multiple Choice Questions from Indian Educational Text
Arunesh Saddish, Pranav Somaiah, Vaasu Gambhir, S. Jaya Nirmala · 2023
Multiple Choice Questions(MCQs) are becoming increasingly popular among testing organizations, educational institutes, and MOOC creators due to their ease of evaluation. With a rapid increase in the number of educational institutes in India, standardized tests in India such as the Joint Entrance Examination(JEE) are adopting MCQs as the format of questioning. Human evaluators find the difficulty of questions generated by existing automatic MCQ generators on the Indian context easy. To avoid this an automatic MCQ generator capable of capturing paragraph level dependencies, framing questions around them, and generating semantically close distractors is proposed, that is optimised on Indian high school geography and biology textbooks of the CBSE curriculum. The MCQ generator has two distinct modules, a question and answer generation module and a distractor generation module. The question and answer generation module contains a T-5 transformer trained on the Stanford Question Answering Dataset(SQuAD) dataset and fine-tuned on a custom handmade dataset to train the model to generate questions linking concepts in different sentences in different parts of the source text and perform well in the context of Central Board of Secondary Education(CBSE) textbooks. The distractor generator uses a combination of the outputs from a Part of Speech(POS) tagger, a T5 transformer, a sense2vec model, and a syntax checker to generate distractor options semantically close to the answer option. The model was judged by human evaluators for fluency, relevance, answerability, grammar, and difficulty and outscored the popular encoder-decoder based question generators. The model also promises a better F1 score than the encoder-decoder based question generators.