Vision–Language Pretraining for Image Captioning Using Facial Expression Recognition

Abdul Saboor Khan, Imran Shafi, Abdul Haseeb Khan · IEEE Access · 2025

This paper presents a novel approach incorporating Facial Expression Recognition (FER) to improve emotional and contextual understanding in Vision-Language Pretraining (VLP) model-generated image captions. Conventional captioning models focus primarily on object identification, producing captions devoid of emotional content. The proposed model integrates FER to generate captions reflecting both visual and emotional characteristics of images. The model employs Vision Transformers (ViTs) for image encoding and BERT for text encoding, trained on datasets including MS COCO, CC3M, SBU Captions, CC12M, and Visual Genome. It is fine-tuned on the More Inclusive Annotations for People (MIAP) dataset to incorporate FER capabilities. The FER-enhanced VLP model outperforms baseline models like BLIP and CLIP across various metrics. The proposed model achieved BLEU-4 score of 41, CIDEr score of 143, and ROUGE score of 61, surpassing baseline performances. Results demonstrate that FER integration significantly improves semantic performance, language quality, and contextual comprehension within the VLP framework. These enhancements highlight the importance of emotional analysis in AI-based captioning services, with applications in assistive technology, content creation, and human-computer interaction. This study shows that FER integration with VLP enhances both vision capabilities and emotional understanding, producing more contextually aware captions.

Read the paper · More papers on PaperTik