Unsupervised Dual Modality Prompt Learning for Facial Expression Recognition

Muhammad Shahid · 2023

A method of facial expression recognition using a vision language model is proposed. Recently vision-language models for example CLIP (Contrastive Language-Image Pre-training) models developed by OpenAI have achieved exceptional results on a variety of image recognition and retrieval tasks, exhibiting strong zero-shot performance. Transferable representations can be adapted through prompt tuning to a variety of downstream tasks. From the general knowledge stored in a pre-trained model, prompt tuning attempts to extract useful information for downstream tasks. In order to avoid time-consuming prompt engineering, recent works use a small amount of labeled data for adapting vision language models to downstream image recognition problems. However, requiring target datasets to be labeled may restrict their scalability. Moreover, we also note that adapting prompt learning techniques in only one branch of CLIP (vision or language) is suboptimal because it won't allow for the dynamic adjustment of both representation spaces on a downstream task. In this paper, we evaluated the performance of the CLIP model as a zero-shot face recognizer and proposed an Unsupervised Dual Modality Prompt Learning framework for Facial Expression Recognition. Our model tunes the prompts through learning text and visual prompts simultaneously to improve alignment between the linguistic and visual representations when labels are not provided for the target dataset. The experimental results on CK+, JAFFE, RAF-DB, and FER2013 datasets showed that our proposed method performs better compared with CLIP Zero-Shot and other unsupervised prompt-based learning methods for facial expression recognition tasks.

Read the paper · More papers on PaperTik