Team HUGE: Image-Text Matching via Hierarchical and Unified Graph Enhancing
Bo Li, You Wu, Zhixin Li · 2024
Graph structures can represent rich semantic relationships, but currently, image-text matching methods have not been well applied. How to efficiently achieve graph learning and prevent overfitting of complex graph-based models are the challenges for all graph-based methods. Besides, single perspective similarity representation learning may overlook some potential correlations. For these issues, we develop Hierarchical and Unified Graph Enhancing (HUGE) to effectively extract the modality variations and collect the corresponding semantics of cross-modal features. Specifically, hierarchical graph learning (HGL) promotes cross-modal learning via training different graph sub-modules from multiple perspectives, while unified graph enhancing (UGE) aims to integrate the plausible alignments from different graph sub-module feedback. Besides, we designed a two-stage similarity representation learning strategy that combines the advantages of cosine similarity and vector similarity. Extensive experiments evidence that our HUGE approach can outperform the SoTA image-text matching methods. Sufficient ablation experiments verify the effectiveness of each component of HUGE.