Temporal Convolutional Networks for Speech and Music Detection in Radio Broadcast

Quentin Lemaire, André Holzapfel · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2019

The task of speech and music detection aims at the automatic annotation of potentially overlapping speech and music segments in audio recordings. This meta-data extraction process has important applications in royalty collection in broadcast audio. This study focuses on deep neural network architectures made to process sequential data, and a series of recent architectures that have not yet been applied for this task are evaluated, extended and compared with a state-of-the-art architecture. Moreover, different training strategies are evaluated, and we demonstrate the advantages of a step-wise procedure that facilitates the combination heterogeneous datasets. The study shows that Temporal Convolution Network (TCN) architectures can outperform state-of-the-art architectures, and that especially the novel extension of non-causal TCN introduced in this paper leads to a significant improvement in the accuracy.

Read the paper · More papers on PaperTik