Prosody Modeling and Eigen-Prosody Analysis for Robust Speaker Recognition

Zi-He Chen, Yuan‐Fu Liao, Yau‐Tarng Juang · 2006

Unseen handset mismatch and limited training/test data are the major source of performance degradation for speaker identification in telecommunication environments. In this paper, a vector quantization (VQ)-based prosody modeling and an eigen-prosody analysis (EPA) is integrated to transform the close-set speaker identification problem into a full text document retrieval-similar task. The prosody modeling labels the prosodic feature contours of a speaker's speech into sequences of prosody states. EPA then constructs a compact eigen-prosody space to represent the constellation of speakers. Furthermore, EPA is fused with a lower-level a priori knowledge interpolation (AKI) handset distortion compensator to complement each other. Experimental results on the HTIMIT database had shown that about 41.0% and 32.8% relative error rate reduction for seen and unseen handsets, respectively, was achieved compared with the maximum a priori-adapted Gaussian mixture model/cepstral mean subtraction (MAP-GMM/CMS) baseline.

Read the paper · More papers on PaperTik