Stacked Attention based Textbook Visual Question Answering with BERT

R Aishwarya, P Sarath, Shibil Rahman P, U Sneha, Sruthy Manmadhan · 2022 IEEE 19th India Council International Conference (INDICON) · 2022

In the process of learning (especially in online learning), students may come across various complex images/diagrams which they might find difficult to understand. They may have doubts and ambiguities regarding its structure, components and usage etc. In such situations, it is difficult to map these questions to the actual facts. In such a scenario students may tend to refer to various websites and textbooks to understand it. This method is quite laborious, time consuming and is found to be less efficient. To overcome these problems, introducing a system that intends to respond to multi-model questions given a textual context, diagrams or images. Students can ask any queries regarding the content of the image and as an outcome, the system will provide a descriptive explanation by combining relevant points from the textbook, as the answer. The proposed system takes an image/diagram and a natural language question from the student as input. The question is analyzed and understood by the system using Bidirectional Encoder Representations from Transformers (BERT). The image is analyzed with computer vision technique and deep learning algorithm like VGG16 using the training dataset built upon specific curriculum-based data. To provide more precise answer these two extracted features are given to Stacked attention network (SAN). The correct answer for the visual question is generated from the system. Since there isn’t any dataset available for the proposed problem, a dataset with textbook diagrams, associated questions and the corresponding answers is created from scratch and used for training. It helps students to develop a curious mind and helps to have a thorough understanding of the concepts.

Read the paper · More papers on PaperTik