Semantic Threads Enabling Image-Text Retrieval via VQA Transformers

Junnubabu Noorbhasha, Rohitha Guddeti, Satvika Lingutla, Sujitha Etikikota, Santhosh Kumkumkari · 2024

The integration of vision and language has propelled the advancement of artificial intelligence systems. Visual Question Answering (VQA) stands at the nexus of computer vision and natural language processing, enabling machines to comprehend and respond to image-related queries. This paper introduces a novel VQA approach harnessing the capabilities of the BLIP (Bootstrapping Language Image Pretrained) model, a transformer-based architecture esteemed for its natural language understanding prowess. The methodology involves image preprocessing, and question translation into a standardized language for efficient processing by BLIP. Mainly, the study integrates multilingual support into the VQA framework, facilitating seamless interaction with users across diverse linguistic backgrounds. Through rigorous experimentation, this paper demonstrates the effectiveness of our approach in accurately answering questions in various languages These findings underscore the robustness and adaptability of the BLIP model in handling multilingual inputs, thereby enhancing accessibility and usability in real-world applications. This research contributes to advancing the state-of-the-art in VQA systems by addressing language barriers and promoting inclusivity in human-machine interaction.

Read the paper · More papers on PaperTik