Triplet Loss with Curriculum Learning for Audio-Visual Retrieval
Donghuo Zeng, Kazushi Ikeda · 2023
The cross-modal retrieval models leverage the potential of triple loss optimization to learn robust embedding spaces. However, existing methods often train these models in a singular pass, overlooking the distinction between semi-hard and hard triples in the optimization process, which will lead to suboptimal model performance. In this paper, we introduce a novel approach rooted in curriculum learning to address this problem. We propose a two-stage training paradigm that guides the model’s learning process from semi-hard to hard triplets. In the first stage, the model is trained with a set of semi-hard triplets, starting from a low-loss base. Subsequently, in the second stage, the model mines the hardest triplet with the primary aim of mitigating the risk of overfitting by addressing the highest loss. Extensive experimental results conducted on the audio-visual dataset show a significant improvement of approximately 9.8% in terms of average Mean Average Precision (MAP) over the current state-of-the-art method, MSNSCA, for the Audio-Visual Cross-Modal Retrieval (AV-CMR) task on the AVE dataset, indicating the effectiveness of our proposed method.