Performance vs. Complexity Comparative Analysis of Multimodal Bilinear Pooling Fusion Approaches for Deep Learning-Based Visual Arabic-Question Answering Systems

Sarah M. kamel, Mai A. Fadel, Lamiaa A. Elrefaei, Shimaa I. Hassan · Computer Modeling in Engineering & Sciences · 2025

Visual question answering (VQA) is a multimodal task, involving a deep understanding of the image scene and the question’s meaning and capturing the relevant correlations between both modalities to infer the appropriate ans... | Find, read and cite all the research you need on Tech Science Press

Read the paper · More papers on PaperTik