Speaker identification in the presence of packet losses

Deva K. Borah, P. DeLeon · 2004

Gaussian mixture model (GMM)-based speaker identification systems have proved remarkably accurate for large populations using reasonable lengths of high-quality test utterances. Test utterances, however, acquired from cellular telephones or over the Internet (VoIP) may have dropouts due to packet loss. In our research, we have demonstrated that for small packet sizes, these losses can result in degraded accuracy of the speaker identification system. It is shown that by training the GMM model with lossy speech packets, corresponding to the loss rate experienced by the speaker to be identified, significant performance improvement is obtained. In order to avoid the prior estimation of the packet loss rate experienced by the test subject, we propose an algorithm to identify the user based on maximizing the a posteriori probability over the GMM models of the users, trained with several packet loss rates. It is shown that the proposed algorithm provides excellent identification performance.

Read the paper · More papers on PaperTik