Speech-music discrimination from MPEG-1 bitstream

Roman Jarina, Noel A. Murphy, Noel Edward O'Connor, Seán Marlow · Arrow@dit (Dublin Institute of Technology) · 2001

Abstract: This paper describes a proposed algorithm for speech/music discrimination, which works on data directly taken from MPEG encoded bitstream thus avoiding the computationally difficult decoding-encoding process. The method is based on thresholding of features derived from the modulation envelope of the frequency-limited audio signal. The discriminator is tested on more than 2 hours of audio data, which contain clean and noisy speech from several speakers and a variety of music content. The discriminator is able to work in real time and despite its simplicity, results are very promising.

Read the paper · More papers on PaperTik