A Gaussian-selection-based preclassifier for speaker identification
Marie A. Roch · The Journal of the Acoustical Society of America · 2002
For practical reasons driven by the need to periodically adapt speaker models or enroll new members of the population, most speaker identification systems train individual models for each speaker. When classifying a speech token, the token is scored against each model and the maximum a posteriori decision rule is used to decide the classification label. Consequently, the cost of classification grows linearly for each token as the population size grows. When considering that the number of tokens to classify is also likely to grow linearly with the population, the total work load increases exponentially. In this work, a new system is presented which builds upon the so-called ‘‘Gaussian selection’’ techniques. The system uses the speaker-specific models as source data and constructs N-best hypotheses of speaker identity. The N-best hypothesis set is then evaluated using individual speaker models. This process results in an overall reduction of workload. The cost of the model generation is low enough to permit enrollment and adaptation, and the accuracy of the preclassifier is such that there is minimal impact on the recognition rate.