FastVDT: Fast Transformer With Optimised Attention Masks and Positional Encoding for Visual Dialogue

Qiangqiang He, Shuwei Qian, Chongjun Wang · IET Computer Vision · 2025

ABSTRACT The visual dialogue task requires computers to comprehend image content and preceding question‐and‐answer history to accurately answer related questions, with each round of dialogue providing the necessary historical context for subsequent interactions. Existing research typically processes multiple questions related to a single image as independent samples, which results in redundant modelling of the images and their captions and substantially increases computational costs. To address the challenges above, we introduce a fast transformer for visual dialogue, termed FastVDT, which utilises novel attention masks and continuous positional encoding. FastVDT models multiple image‐related questions as an integrated entity, accurately processing prior conversation history in each dialogue round while predicting answers to multiple questions. Our method effectively captures the interrelations among questions and significantly reduces computational overhead. Experimental results demonstrate that our method delivers outstanding performance on the VisDial v0.9 and v1.0 datasets. FastVDT achieves comparable performance to VD‐BERT and VU‐BERT while reducing computational costs by 80% and 56%, respectively.

Read the paper · More papers on PaperTik