Neural network models for combining evidence from spectral and suprasegmental features for text-dependent speaker verification

S. R. Mahadeva Prasanna, J.M. Zachariah, B. Yegnanarayana · 2004

This paper proposes a method using neural network models for combining evidence from spectral and suprasegmental features for text-dependent speaker verification. Spectral features are extracted using the Dynamic Time Warping (DTW) technique. While extracting the spectral features, the DTW algorithm is used only to obtain a matching score and the information present in the warping path is ignored. In this work a method is discussed to extract suprasegmental features such as pitch and duration using the information in the warping path. Although the suprasegmental features may not yield good performance, combining the evidence from suprasegmental and spectral features improves the performance of the speaker verification system significantly.

Read the paper · More papers on PaperTik