Algorithm to detect the beginning and end points of a speech utterance

K. Ganesan, Wen‐Chen Lin · The Journal of the Acoustical Society of America · 1975

There is a great need to detect the beginning and end points of a speech utterance in applications like speech recognition and speaker identification. In this paper, we present a method for beginning and end-point detection which makes use of the maximum likelihood principle. The features that are used by the algorithm are (1) total per-unit energy, (2) zero-crossing rate, and (3) absolute amplitude of the speech samples, Conditional probability densities are estimated for these three features using a database of 60 phonetically balanced words and ten phonetically balanced sentences spoken by four male speakers with General American accents. A set of optimum thresholds are obtained for each feature such that the probability of classification error is minimized. The algorithm was tested for both isolated words and sentences over a population of six speakers and an error rate of nearly 0% was observed.

Read the paper · More papers on PaperTik