Beyond Words: Exploring Co-Attention with BERT in Visual Question Answering

Vibhashree B. S, Nisarga Kamble, Sagarika Karamadi, Sneha Varur, Padmashree Desai · 2024

The Visual Question Answering project described in this work uses a multi-modal approach that combines Bidirectional Encoder Representations from Transformers for complex natural language question understanding, You Only Look Once for robust object detection, and a Hierarchical Co-Attention Mechanism for efficient integration of textual and visual data. The goal is to develop a thorough Visual Question Answering system that can respond to inquiries about visual information with precision and contextual relevance. While Bidirectional Encoder Representations from Transformers captures subtle question semantics, the You Only Look Once model makes efficient object detection possible. In order to handle the complexities of multi-modal information processing, the Hierarchical Co-Attention Mechanism improves the fusion of local and global characteristics. The evaluation criteria include co-attention fusion scores, object detection metrics, Natural Language Processing metrics for Bidirectional Encoder Representations from Transformers, and overall Visual Question Answering accuracy, among other pertinent dimensions. By taking into account both response correctness and the efficiency of individual components, the study seeks to obtain a sophisticated view of the system's performance. Open-ended evaluation criteria allow for a comprehensive assessment because of the Visual Question Answering assignment's varied nature. The results show that the system can intelligently answer to questions regarding visual content and demonstrate its potential for a wide range of practical applications. The project is open-sourced, encouraging community involvement and collaboration to push the boundaries of multi-modal question answering systems even farther.

Read the paper · More papers on PaperTik