Text-dependent speaker verification using subword neural tree networks
Han-Sheng Liou, Richard J. Mammone · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1994
In this paper, a new algorithm for text-dependent speaker verification is presented. The algorithm uses a set of concatenated Neural Tree Networks (NTNs) trained with sub-word units for speaker verification. The conventional NTN when trained by all the words in training data achieves good results in the text-independent task. The proposed method is described as follows. First, the predetermined password in the training data is segmented into sub-word units by Hidden Markov Model (HMM). Second, an NTN is trained for only the data segmented into that sub-word unit. It integrates the discriminatory ability of NTN with the framework of HMMs which demonstrates ability in modeling temporal variation of speech. The sub-word NTN trained with clustered data reduces the complexity of the NTN structure, and is more powerful in discriminating speakers. This new algorithm was evaluated by experiments on a TI isolated-word database, which contains 16 speakers. An improvement of performance was obtained over baseline performance obtained from conventional methods.