Unsupervised speaker segmentation in telephone conversations
A. Cohen, V. Lapidus · 2002
Speaker recognition and verification has been used in a variety of commercial, forensic and military applications. The classical problem is that of supervised recognition, in which there is sufficient a priori information on the speakers to be identified. This paper deals with the problem of unsupervised speech segmentation and speaker classification, where no a priori speaker information is available. The algorithm accepts dual-speaker conversation telephone speech data, detects events of simultaneous speakers, and segment the signal by assigning each speech segment to its speaker. Discrete HMM are used, with 12th order cepstral coefficients. Correct recognition rates of more than 90% are demonstrated.