On-Line Speech/Music Segmentation for Broadcast News Domain
Marko Kos, Matej Grašič, Damjan Vlaj, Zdravko Kačič · 2009
This paper presents novel feature-group for on-line speech/music segmentation for broadcast news domain. The features are based on mel-frequency cepstral coefficients variance (MFCCV). The idea behind the feature-group construction is the energy variation in a narrow frequency sub-band. The variation is bigger for speech than for music. For feature discrimination and segmentation ability evaluation the radio broadcast database was used. Results show that MFCCV features perform better than the classic MFCC features. The MFCCV features are very convenient speech/music discriminator for automatic speech recognition system where MFCC features are used, as they perform better than classic MFCC features and only one additional calculation step is needed.