Gmm-based speaker recognition for mobile embedded systems
Cheung-Chi Leung, Y.S. Moon · 2004
When porting speaker recognition systems originally targeted for desktop personal computers to mobile embedded systems, the most often encountered problem is unsatisfactory performance in terms of execution time. This thesis addresses this performance issue in four directions. Specifically, our speaker recognition system uses mel-frequency cepstral coefficient feature extraction and GMM-based pattern matching techniques. Firstly, we study the effect of window size and shift period in mel-frequency cepstral coefficient feature extraction on computation time and accuracy in a typical speaker verification system. Experiments show that a critical point exists in choosing a window size in feature extraction so that using a window size less than the critical point results in tradeoffs between accuracy and computation time. Secondly, we adopt a number of fixed-point arithmetic optimization techniques, including range estimation, floating-point code conversion and transcendental function optimization, to enable real-time speaker verification in mobile embedded systems. Our work shows that the optimization techniques can provide an execution-time speedup factor up to 37 times over a baseline system, while the verification accuracy is not affected. Thirdly, we propose a pruning algorithm which attempts to make the verification decision by considering only a portion of the testing utterance. With appropriate threshold values learnt from a set of training data, the pruning approach could reduce the execution time of pattern matching on the entire testing set by 20% while maintaining the verification performance. We also formulate a tree-based speaker modeling approach to minimize the execution time in speaker identification in the fourth part. The proposed algorithm attempts to reduce the computational complexity in the identification phase from K to log K (where K is the number of registered speakers) by collecting extra information (similarity among registered speakers) during the enrollment phrase. Experiments with 60 speakers show that the tree-based approach can provide a computation-time speedup factor of 2, while reducing the identification accuracy by 3% only. This thesis illustrates that although the computation power of the mobile embedded devices is limited, real-time speaker recognition can still be achieved in such devices by adopting the previously mentioned optimization techniques, with nearly the same recognition accuracy as performed in most desktop personal computers.