Exploring the impact of advanced front-end processing on NIST speaker recognition microphone tasks.
William M. Campbell, Douglas E. Sturim, Bengt Borgström, Robert B. Dunn, Alan V. McCree, Thomas F. Quatieri, Douglas A. Reynolds · 2012
The NIST speaker recognition evaluation (SRE) featured mi-crophone data in the 2005-2010 evaluations. The preprocess-ing and use of this data has typically been performed with tele-phone bandwidth and quantization. Although this approach is viable, it ignores the richer properties of the microphone data— multiple channels, high-rate sampling, linear encoding, ambient noise properties, etc. In this paper, we explore alternate choices of preprocessing and examine their effects on speaker recogni-tion performance. Specifically, we consider the effects of quan-tization, sampling rate, enhancment, and two-channel speech activity detection. Experiments on the NIST 2010 SRE inter-view microphone corpus demonstrate that performance can be dramatically improved with a different preprocessing chain. 1.