Advancing Multimodal Classification: An In-depth Analysis of Amazon Reviews Using Early Fusion with Gating Mechanism
Aritra Mandai, Wazib Ansar, Amlan Chakrabarti · 2024
In today’s digital landscape, considering diverse modalities of data has become vital for holistic analysis of customer reviews often containing both text and images. Traditional approaches focus mainly on text, neglecting valuable visual information. Our paper proposes a novel early fusionbased gating mechanism for multimodal classification of Amazon reviews. For generation of textual features BERT has been utilized whereas visual features have been obtained through ResNet. By leveraging gated integration of textual and visual information while enforcing alignment among modalities using Contrastive Language-Image Pre-Training (CLIP), our model offers a comprehensive understanding of consumer opinions. This approach is particularly useful for identifying inconsistencies, verifying product authenticity, and enhancing review-based insights’ reliability. Our model excels the SOTA approaches with a commendable F1 score of ${0. 8 8}$.