The optimization of perceptually-based features for speaker identification
Xu Li, John Oglesby, J.S. Mason · International Conference on Acoustics, Speech, and Signal Processing · 2003
Results of an experimental study and the optimization of features for a conventional vector-quantization codebook-based automatic speaker identification (ASI) system are presented. Standard LPC (linear predictive coding) and a perceptually weighted feature termed PLP (perceptually based linear prediction) are compared using appropriate distance measures, namely, the log-likelihood, and three cepstral variants: constant weighting, the robot-power-sum, and the inverse variance. PLP features combined with a weighted cepstral measure are found to be consistently the best in a number of different digit-independent ASI experiments. Results support the hypothesis that the higher orders of PLP (>5) contain significant speaker-specific information, with ASI performance improving rapidly up to order 8, and then far more slowly yet consistently up to order 16. A similar pattern is seen for codebook size, with fast improvements up to size 64, with more gradual gains thereafter.>