Story segmentation and topic detection for recognized speech
S. Dharanipragada, Martin Franz, J. S. McCarley, Salim Roukos, T. Ward · 1999
We present a technique for the segmention of a sound track into two classes of segments. Each frame of signal is preprocessed by extracting cepstral coefficients and their first order derivatives. For each class, the distribution of the frame parameter vectors is modeled by a Gaussian Mixture Model (GMM). GMM order is selected using two criteria : the Minimum Description Length (MDL) criterion and the Akaike Information Criterion (AIC). Frame score is based on a weighted loglikelihood ratio in a window around the frame. Decision for each frame is taken by comparing its score to a threshold. Experiments are presented on speech / music segmentation in audio tracks. In these experiments, the MDL criterion leads to a reasonable GMM order. Using the MDL criterion for GMM order selection, frame classification error rate is around 20%. However, using GMMs with much lower orders, only decreases marginally performances.