Evaluating Vision Transformers Efficiency in Image Captioning

Bintang Kevin Hizkia Samosir, Sani Muhamad Isa · 2024

This study investigates the performance of Vision Transformer (ViT) variants—the Shifted Window Transformers (SWIN), Distillation with No Labels (DINO), and Data-efficient Image Transformers (DeIT)—in image captioning tasks using the Flickr8K dataset. While ViT architectures have shown promise in image classification, their effectiveness for image captioning, particularly with smaller datasets, remains unclear. The models' performance was evaluated using BLEU metrics, while training efficiency was analyzed through Pareto front analysis of computational time and accuracy. Among the tested variants, SWIN Transformers demonstrated superior performance (BLEU-1: 64.4, BLEU-2: 33.9, BLEU-3: 17.1, BLEU-4: 8.4), followed by DINO (BLEU-1: 63.1, BLEU-2: 32.7, BLEU-3: 16.4, BLEU-4: 7.5), while DeIT showed the weakest performance (BLEU-1: 61.6, BLEU-2: 31.1, BLEU-3: 14.7, BLEU-4: 6.5). SWIN Transformers achieved the shortest training time at 3 minutes 31 seconds per epoch, making it the most efficient model among ViT variants based on Pareto front analysis. While ViT variants achieved competitive BLEU-1 scores comparable to previous top models, they struggled with generating coherent, longer sentences, as evidenced by suboptimal BLEU-4 scores. These findings provide empirical evidence of how the lack of inductive bias in transformer architectures affects their ability to capture complex scene relationships, despite their strong feature detection capabilities, contributing to the understanding of transformer models' limitations in vision-language tasks, especially with limited data.

Read the paper · More papers on PaperTik