Modal Feature Contribution Distribution Strategy in Visual Question Answering

Feng Dong, Xiaofeng Wang, Ammar Oad, Masood Nazir Khoso · Journal of Engineering Science and Technology Review · 2022

The results of the visual question answering (VQA) task are obtained by the joint inference of the information of the image and text modalities, and its performance is affected by multiple factors, such as the modal feature extraction method and the modal feature fusion process.The current popular VQA models do not undertake processing after extracting the modal features from images and texts.Instead, feature fusion is executed straight without considering the modality's feature contribution, which is debatable.To reveal the relationship between the contribution distribution of image and text modal features, and the performance of VQA task, a plug-and-play contribution distribution strategy of modal features was proposed based on nonlinear functions.By globally operating the image and text features on the basis of extracting the features of the two modalities, the feature weight distribution in each mode was processed and analyzed based on the nonlinear function and the global features of the two modes.Extensive experiments were carried out on the basis of the existing model to fully verify the effectiveness of this strategy.Results show that the contribution distribution of the extracted image and text features before the feature fusion stage of the modality harmonizes the relationship between the image and text modalities and strengthens the effective features in the respective modalities.At the same time, the strategy further improves the performance of existing VQA models.This study provides a certain reference on how to reprocess features in VQA tasks.

Read the paper · More papers on PaperTik