Sentiment Recognition in Images leveraging ResNet18 vs Vit Architecture

Paladugu Trisha Sai, Gudavalli Hruthi Sri, T.Lakshmi Surekha · 2024

In the field of natural language processing (NLP), the analysis of sentiment detection in photos is crucial since it makes it easier to decipher and comprehend the feelings of the people shown in the images. by carefully contrasting Vision Transformer (ViT) designs with Residual Network (ResN et18) topologies in Deep Learning to evaluate their effectiveness in this situation. While RESNET18s have long been the cornerstone of image processing, ViT is a promising upstart that has demonstrated exceptional performance in a variety of computer vision workloads. The intention is to use these models for the job of face emotion recognition in humans. In this study, RESNET18 and ViT models that have been trained on a variety of datasets of photos with sentiment labels attached to human faces are being developed and implemented. Strong representations of human facial expressions may be learned by the models thanks to the training dataset's wide range of emotions. By utilizing cutting-edge techniques to guarantee peak performance during training and carrying out comprehensive tests to assess precision, effectiveness, and robustness in sentiment analysis tasks for both architectures. A full comparison study is also provided by mentioning the model complexity, processing needs, scalability, and graphic rendering methodologies.

Read the paper · More papers on PaperTik