Prediction-error-based Adaptive SpecAugment for Fine-tuning the Masked Model on Audio Classification Tasks

Xiao Zhang, Haoran Xing, Mingxue Song, Daiki Takeuchi, Noboru Harada, Shoji Makino · 2024

Spectrogram augmentation (SpecAugment), a data processing method that enlarges datasets without adding new samples, has been widely employed in the fine-tuning process of the masked model for audio classification. In the conventional SpecAugment, the positions of time masking and frequency masking, which directly determine the available information that the model can learn from, are selected randomly. As a result, the random position masking in the conventional SpecAugment may prevent the model from fully utilizing the input information. To this end, we propose a Prediction-error-based Adaptive SpecAugment (PEAS), which incorporates two auxiliary tasks based on reconstruction and then introduces a mask position selector in the fine-tuning process for the masked model. Rather than masking at random positions, the proposed PEAS generates masks on the parts of the spectrograms that are hard to reconstruct for the model. Masking these positions forces the model to learn more generalized audio features, which can effectively prevent the model from learning classification labels by identifying individual special features. Besides, to accelerate the learning of audio features during the early training epochs, we progressively increase the proportion of adaptive masks. Experimental results demonstrate that our proposed PEAS can match or outperform both the conventional method and random masking strategy on the ESC-50 and Speech Commands V2 datasets.

Read the paper · More papers on PaperTik