Histogram transform model using MFCC features for text-independent speaker identification
Hong Tao Yu, Zhanyu Ma, Minyue Li, Jun Hai Guo · 2014 48th Asilomar Conference on Signals, Systems and Computers · 2014
A novel text-independent speaker identification (SI) method is proposed in this paper. This method uses the mel-frequency cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature set to capture the speaker's characteristics. In order to utilize dynamic information, we design super MFCCs feature by cascading 3 neighboring MFCCs frames together. The probability density function (PDF) of these super MFCCs features is estimated by the recently proposed histogram transform (HT) method, which generated more training data by random transforms to realize the histogram PDF estimation and recede the discontinuity problem of the common multivariate histograms computing. Compared to the conventional PDF estimation method, such as Gaussian mixture model, the HT model shows promising improvement in a SI task.