FeatureCut: An Adaptive Data Augmentation for Automated Audio Captioning

Zhongjie Ye, Yuqing Wang, Helin Wang, Dongchao Yang, Yuexian Zou · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

Automated audio captioning (AAC) aims at generating natural language descriptions for an audio clip. Due to the difficulty and high cost of annotating audio-caption pairs, the existing audio captioning dataset is of a very small scale which leads to unsatisfied performance for AAC models. One intuitive and effective solution is to augment training data to boost performance instead of annotating more data. To this end, we propose an online data augmentation method (FeatureCut) incorporating the encoder-decoder framework to enable the language decoder fully make use of the acoustic features in generating the captions. Specifically, we propose to construct a Feature Attention-weight Table (FAT) from the attention-weight maps given by the attention module in the language decoder for each audio clip. That is, the FAT can reflect the importance of each acoustic feature in the decoding process. Then a progressive acoustic features cutting strategy is designed to discard the higher magnitude values in FAT to generate augmented acoustic features. We apply Kullback-Leibler divergence (K-L divergence) between original and augmented data to encourage AAC models to make similar predictions from different views of them, in order to balance the learning capability of ACC models. We evaluate our proposed FeatureCut on the Clotho dataset. The experimental results demonstrate that our proposed FeatureCut can be easily plugged into the state-of-the-art ACC models (i.e. MAAC and CNN-Transformer) and can obtain 5.1% and 3.1% gains in CIDEr scores.

Read the paper · More papers on PaperTik