An entropy based robust speech boundary detection algorithm for realistic noisy environments

K. Weaver, Khurram Waheed, F.M. Salem · 2004

This paper addresses the issue of automatic word/sentence boundary detection in both noiseless and noisy backgrounds. We present our proposed speech boundary detection algorithm using a time-domain entropic contrast function. The entropic contrast exhibits well-behaved characteristics as compared to energy-based methods resulting in immunity to endpoint cut-of issues for the latter. This algorithm is capable of estimating the speech boundaries both in noiseless and distinctive noise backgrounds such as a fan, car engine, radio etc. For the case of wide-spectrum colored background noise such as jazz, opera, songs, rock music etc., we further propose a modification in the preprocessing stage by incorporating a frequency-weighting scheme to emphasize the speech contents. This improved scheme provides proper speech segmentation even in the presence of wide-spectral background noise with no change in the computational cost versus our earlier proposed algorithm. A complete time-domain implementation is sought due to its lower computational burden and its suitability for real-time implementations using DSPs, FPGAs, ASICs etc. The algorithm improves the accuracy of word boundary estimates by a factor of at least 25% for the case of isolated (and 16% for connected) speech. For continuous speech, the algorithm can determine sentence boundaries thus allowing for power efficient implementation of speech recognition engines by rejecting extended periods of silence.

Read the paper · More papers on PaperTik