An entropy based robust speech boundary detection algorithm for realistic noisy environments
K. Weaver, Khurram Waheed, F.M. Salem · 2004
This paper addresses the issue of automatic word/sentence boundary detection in both noiseless and noisy backgrounds. We present our proposed speech boundary detection algorithm using a time-domain entropic contrast function. The entropic contrast exhibits well-behaved characteristics as compared to energy-based methods resulting in immunity to endpoint cut-of issues for the latter. This algorithm is capable of estimating the speech boundaries both in noiseless and distinctive noise backgrounds such as a fan, car engine, radio etc. For the case of wide-spectrum colored background noise such as jazz, opera, songs, rock music etc., we further propose a modification in the preprocessing stage by incorporating a frequency-weighting scheme to emphasize the speech contents. This improved scheme provides proper speech segmentation even in the presence of wide-spectral background noise with no change in the computational cost versus our earlier proposed algorithm. A complete time-domain implementation is sought due to its lower computational burden and its suitability for real-time implementations using DSPs, FPGAs, ASICs etc. The algorithm improves the accuracy of word boundary estimates by a factor of at least 25% for the case of isolated (and 16% for connected) speech. For continuous speech, the algorithm can determine sentence boundaries thus allowing for power efficient implementation of speech recognition engines by rejecting extended periods of silence.