Efficient Visual Question Answering on Embedded Devices: Cross-Modality Attention With Evolutionary Quantization

Aakansha Mishra, Aditya Agarwala, Utsav Tiwari, Vikram Nelvoy Rajendiran, Srinivas Soumitri Miriyala · 2024

Visual Question Answering (VQA) lies at the intersection of vision and language domains necessitating learning representations from multiple modalities. While the model development for VQA has witnessed tremendous growth, the efforts for its deployment on embedded devices have been lagging limiting its true potential. In this work, the authors address this challenge by designing a novel hardware-friendly architecture for VQA based on the transformer model with cross-modality attention. The memory footprint of the VQA model is optimized for on-device deployment using a distributed framework for Post Training Quantization (PTQ) formulated as a Non-Linear Programming (NLP) problem. The NLP problem is solved using an Evolutionary algorithm to determine the low-bit representation of the VQA model with minimal accuracy drop compared to the full precision model. The quantized model for VQA with a marginal accuracy drop of less than 2%, resulted in 4 times memory improvement, and over 2 times latency improvement, enabling its successful deployment on the Samsung Galaxy S23 device. The comprehensive study explores the potential of the proposed generic end-to-end pipeline from VQA model development to its deployment.

Read the paper · More papers on PaperTik