Transfer learning based on self-attention mechanism in sound event detection

Zihan Liu, Xichang Cai, Jiaxin Li, Ziling Qiao · 2022 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology (CEI) · 2022

SED is a task to detect sound segments. Recent research has found that models based on self-attention mechanism can have an edge on sound feature detection. However, the self-attention mechanism requests a large-scale data set which is difficult to achieve in the audio field. So, this paper proposes a transfer learning-based SED solution. Through the transfer of AST based on the self-attention mechanism, our model reduces the demand for the data. At the same time, we propose a new audio embedding processing method and significantly shorten the training time. After training using the DESED data set on the DCASE task, it completely ahead of CRNN. It reaches the Intersection-based F1 score of 91.4%, and 42.1% higher than CRNN, which reflects the superiority of the self-attention mechanism.

Read the paper · More papers on PaperTik