3CMLF: Three-Stage Curriculum-Based Mutual Learning Framework for Audio-Text Retrieval

Yi-Wen Chao, Dongchao Yang, Rongzhi Gu, Yuexian Zou · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

Audio-text retrieval aims to retrieve instances that best match a given instance from an audio modality to a text modality and vice versa. Recent studies have mainly focused on capturing the shared high-level semantic concepts between these two modalities by synchronously updating the audio and text encoders. We found that such a synchronous updating strategy results in sub-optimal learned audio and text encoders owing to the two encoders' varying initial prior knowledge level. Furthermore, we observed a big semantic gap between the representation of audio and text encoders using the common mini-batch sampling strategy. To tackle these issues, we present a novel three-stage curriculum-based mutual learning framework (3CMLF) to boost the performance. Our approach includes two key components: (i) Inspired by the human learning process, we provide a global curriculum-based hard sample mining strategy, which can globally mine the easiest, median, and hardest negative samples from the full training set and construct three training sets respectively. (ii) We propose to train the text and audio encoders under the three-stage cross-modal mutual learning framework using the three constructed training sets. In the first stage, we fix the weights of the text network, which are initialized using a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model, and then update the audio encoder based on the easiest training set. During the second stage, we freeze the audio encoder and update the text network based on the median training set. After these initial alignment stages, we release all weights to be learned and fine-tuned on the hardest training set. This three-stage process is crucial for allowing the model to successfully differentiate the top retrieved instance from a hard negative set and capture the correlation between the audio-text modal. Notably, 3CMLF is adaptable to the majority of current audio-text models as it requires no alteration to the model architecture. Experimental results on the AudioCaps dataset show that our method achieves a new state-of-the-art performance.

Read the paper · More papers on PaperTik