Speaker recognition model using two-dimensional mel-cepstrum and predictive neural network
Tadashi Kitamura, Shota Takei · 2002
Describes a speaker recognition model using a two-dimensional mel-cepstrum (TDMC) and a predictive neural network. The speaker model consists of two networks. The first one is a self-organizing vector quantization (VQ) map (Kohonen feature map). The second one is a predictive network, and it learns transitional patterns on the feature map of each speaker's model. The TDMC consists of the averaged features and the dynamic features of the 2D mel-log spectra in the analyzed interval. The measure for speaker recognition is obtained by using a combination of the VQ distortion on the feature map and the prediction error on the predictive network. In this study, text-independent speaker identification experiments for eight speakers were carried out. The experimental results have shown that a combination of a feature map and a predictive network is very effective, and that the proposed model using a TDMC shows robustness for the time interval.