Long Recording Segmentation Based on Simple Power Voice Activity Detection with Adaptive Threshold and Post-Processing

Josef Rajnoha · 2009

This paper describes the method of long recording segmentation based on Voice Activity Detection (VAD). Power based detection using an adaptive threshold derived from power dynamics is the core of presented approach. Simple post-processing based on long time sub-segmentation is used for smoothing of primary VAD output to obtain target start-point and end-point detection of particular utterances within long recordings. Because the algorithm is based on simple power VAD it can be much more easily implemented in comparison to approaches based on speech recognition. Though presented approach is so simple it gives quite robust and satisfactory results for pure segmentation task. The tests with two different data types proved satisfactory results same as practical usage during the creation of new speech corpora.

Read the paper · More papers on PaperTik