Performance Evaluation of SNR Estimation Methods in Forensic Speaker Recognition
Francesco Beritelli, S. Casale, Rosario Grasso, Andrea Spadaccini · 2010
Speech signal quality is of fundamental importance for accurate speaker identification. The reliability of a speech biometry system, in fact, is known to depend on the amount of material available, in particular on the number of vowels present in the sequence being analysed and on the quality of the signal. This paper highlights the performance of two Signal-to-Noise Ratio (SNR) estimation methods (manual and semi-automatic) usually adopted for the evaluation of speech signal quality. The results not only demonstrate the different impact of noise on the single vowels, but also show that often the SNR is over-estimated or under-estimated, and this leads respectively to the inclusion of bad quality biometric samples or the exclusion of good quality data, with a negative impact on the accuracy of the identity verification test. The paper proposes a series of issues that need to be tackled in order to develop a better procedure for the selection of biometric samples extracted from the intercepted audio signal.