A Lightweight Multi-Grained Image-Text Retrieval Paradigm via Cascaded Representation Learning and Parameter-Free Feature Aggregation
Chenyu Lu, Nan Zhang, Shiliang Sun · IEEE Transactions on Circuits and Systems for Video Technology · 2024
Multi-grained cross-modal image-text retrieval models have demonstrated promising outcomes through the alignment of local and global features. However, this advancement often results in larger model sizes and higher computational requirements, which raises concerns regarding the balance between performance and efficiency. To address this challenge, we introduce a novel lightweight multi-grained (LMG) image-text retrieval paradigm aimed at enhancing model efficiency. Specifically, in our approach, we first re-frame the retrieval problem as a cascaded representation learning task. This involves leveraging only fine-grained features to capture coarse-grained constraints, thereby reducing computational burden while maintaining accuracy. Furthermore, we replace computationally expensive parametric feature aggregation methods with three efficient parameter-free alternatives: auto-correlation matrix, discrete linear convolution, and discrete Fourier transform. The proposed LMG model is extensively compared with state-of-the-art approaches on two benchmark datasets, i.e., Flickr30K and MSCOCO, and the experimental results highlight the superior performance of LMG. Additionally, we explore the impact of different feature aggregation methods on LMG and conduct a sensitivity analysis on the coarse and fine-grained constraints ratio hyper-parameter.