MobileViT architecture for Facial Beauty Prediction
Djamel Eddine Boukhari, Ali Chemsa, Zine-Eddine Baarir · 2024
Facial beauty prediction is a complex task that involves quantifying human facial attractiveness based on visual features. In this paper, we propose a novel approach utilizing the MobileViT architecture, a hybrid model that integrates convolutional neural networks (CNNs) and vision transformers (ViTs) to tackle this challenge. The MobileViT model is specifically designed to capture both local details and global context, making it highly effective for facial image analysis while ensuring computational efficiency. We evaluated the performance of our model on the SCUT-FBP5500 dataset, a large-scale facial beauty dataset comprising 5,500 images labeled with beauty scores. Our approach achieves state-of-the-art results, demonstrating a Pearson correlation of 0.9515. These results significantly surpass those of existing methods, including CNN-based and semi-supervised approaches. The lightweight and efficient design of MobileViT renders it suitable for real-time applications on mobile and edge devices, highlighting its potential for practical use in beauty prediction tasks.