Enhancing Multimodal Data Fusion for Fine-Grained Product Classification
Wiliyanti, Nanik Suciati · 2025
This research explores the performance of finegrained classification in retail product leaflets with various multimodal data fusion techniques. We use ResNet50 for image feature extraction and DistilBERT for processing offers and discount text, using an ensemble learning approach to combine the output of both models. Experiments were conducted on the Retail Product Classification 2023 dataset, which includes product images along with advertisement descriptions. Various fusion strategies, early, intermediate, and late fusion, were evaluated to determine the most effective integration approach. The experimental results show that intermediate fusion outperforms the other strategies, achieving 97.02% classification accuracy and 94.7% F1 score. These findings confirm that the intermediate fusion technique significantly improves the performance of multimodal classification models compared to early and late fusion. The results of this study contribute to the development of a more accurate and efficient system for analyzing multimodal data fusion in computer vision techniques.