Speaker and text independent language identification using predictive error histogram vectors
Qian-Rong Gu, T. Shibata · 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03). · 2003
A predictive vector quantization (Gray, M., 1984; Jain, A.K. et al., 2000) based speaker and text independent language identification system is proposed, which uses the statistical distribution of predictive error vectors to recognize the language spoken by native speakers. According to Stan C. Kwasny et al. (see Proc. 5th Midwest Artificial Intelligence and Cognitive Science Soc. Conf., p.53-7, 1993), most high level features of speech, such as tone of voice, rhythm, style, pace, accent, etc., appear to be related to distributional patterns or statistical aggregates of speech waveforms. We further develop the method used by Qian-Rong Gu and Tadashi Shibata (6th World Multiconference on Systemics, Cybernetics and Informatics - SCI2002, 2002) to extract these statistical distributional patterns directly from raw speech waveforms and then use them to identify language. The system has been trained and. tested by speech from English and Japanese native speakers. A best identification ratio of 76.8% can be achieved by our system.