Natural Language Processing based Visual Question Answering Efficient: an EfficientDet Approach

Rahul Kumar Gupta, Parikshit Hooda, Sanjeev, Nikhil Kumar Chikkara · 2020

This research paper is intended to propose a technique to solve the task of image-based question answering using EfficientDet and Bidirectional LSTM (BiLSTM). Visual Question Answering (VQA) requires a combination of Image recognition and Natural Language Processing (NLP) techniques. The proposed technique uses EfficientDet for image processing and BiLSTM for question processing. The model acts efficiently because of efficient image processing. The model takes an image and a question as input processes them independently, then fuses the processed versions and predicts the answer to the question. It outputs the result in an open domain based on input questions on the image. This study intends to help the visually impaired person. The model achieves similar performance as its contemporary models on the benchmark VQA dataset.

Read the paper · More papers on PaperTik