Performance and Cost Balancing in Vision Transformer-Based Image Captioning
Yan Lyu, Yong Liu, Qiangfu Zhao · 2023
Image captioning connects computer vision and natural language processing. Many deep learning models have been proposed for solving this problem, such as Transformer-based, CLIP-based and Diffusion-based models. However, the primary focus has been on increasing the accuracy of generating human-like descriptions for given images, leading to expensive SOTA models that cannot be implemented on computation-limited devices. For better performance, we propose a ViT-LSTM model for solving image captioning tasks by addressing the challenge of long-range dependencies. Our model consists of a ViT model pre-trained using the ImageNet 21k dataset, which captures global context, and an LSTM that generates captions reflecting both local and global visual cues. Additionally, we use convolutional layers to reduce feature map dimensionality while preserving spatial relationships and local patterns. This allows the model to understand local details and relationships between regions, achieving better performance with less computational cost.