Multimodal Deep Learning: Integrating Text and Image Embeddings with Attention Mechanism
Arpit Bansal, Krishan Kumar, Sidharth Singla · 2024
People make a lot of decisions by following different modalities. In today’s era, the classification problems that we are dealing with are also multimodal. This field has gained significant attention due to the growth of multimedia data on the internet where textual information is seen occasionally and independently associated with images, sounds, or videos. Utilizing transformers has shown effective results. However, it is still unclear to establish the relationship between various modalities. This paper employs CLIP as the foundational model for multimodal classification. The embeddings of image and text are extracted from their respective encoders. The use of the Attention mechanism along with the CLIP Text Embeddings presents the best results on the datasets: UPMC Food-101 and IndoFashion with an accuracy of $96.47 \%$ and $96.87 \%$.