Minimum Classification Error Interactive Training for Speaker Identification
Y. Kida, Hiroki Yamamoto, Chiyomi Miyajima, K. Tokuda, Tadashi Kitamura · 2006
This paper describes an online discriminative training algorithm aiming at achieving speaker identification on interactive robots. A robot incrementally acquires speakers' voice characteristics during the interaction with the speakers. We simulate the situation that the speakers never give their IDs and the robot can only know whether the identification decision was correct or not from the speaker's positive or negative behavioral reaction. The speaker models are adjusted based on this limited information using minimum classification error (MCE) training consisting of positive and negative adaptation. In cases of correct identification, the conventional MCE training algorithm can be used. We compare three kinds of negative adaptation algorithms for the cases of incorrect identification. Experimental results show that the combination of the positive and negative adaptation achieves faster convergence, and negative adaptation which adjusts only a misclassified speaker model reaches an identification rate of 80% four times faster than the positive adaptation alone.