Cepstral mean based speech source discrimination

Adam Greenhall, Les Atlas · 2010

This paper presents and compares methods for discrimination between speech from a broadcast audio device - like a television, radio, or GPS receiver - and live speech in the same acoustic environment. A solution to this discrimination problem has direct application wherever the audio from such a device interferes with voice recognition, verification, or transcription tasks. The methods and theory applied also have potential applications in multimedia and speaker segmentation, as well as in speaker verification. This paper presents a new use of the cepstral mean as an estimator of the linear time-invariant response of a “speaker” - either broadcast or live - over a relatively long time window. The problem is framed in terms of traditional speaker verification, but with two classes of speakers. This method is tested on five different data sets and the results compared for different feature sets, training methods, and window lengths.

Read the paper · More papers on PaperTik