Hybrid Speaker-Based Segmentation System Using Model-Level Clustering

Hyoung‐Gook Kim, D. Ertelt, T. Sikora · 2006

We present a hybrid speaker-based segmentation, which combines metric-based and model-based techniques. Without a priori information about the number of speakers and speaker identities, the speech stream is segmented in three stages: (1) the most likely speaker changes are detected; (2) to group segments of identical speakers, a two-level clustering algorithm is performed using a Bayesian information criterion (BIC) and HMM model scores - every cluster is assumed to contain only one speaker; (3) the speaker models are reestimated from each cluster by HMM. Finally a resegmentation step performs a more refined segmentation using these speaker models. To measure the performance, we compare the segmentation results of the proposed hybrid method versus metric-based segmentation. Results show that the hybrid approach using two-level clustering significantly outperforms direct metric-based segmentation.

Read the paper · More papers on PaperTik