A voice activity detection algorithm and comfort noise for communication systems with dynamically varying background acoustic noise
Ick Don Lee, Harold P. Stern · 1997
Speech can be modeled as short bursts of vocal energy (talkspurts) separated by gaps containing no vocal energy (silence gaps). Silence gaps occur between syllables and between words, as well as between phrases and between sentences. During a typical conversation each party generates talkspurts approximately 31.5% of the time and is silent the remaining 68.5% of the time. Significant spectral efficiency can thus be achieved by disconnecting the user from the spectral resource during the silent periods and allowing other users access to the spectrum. This discontinuous transmission also reduces average transmitted power from the user, thus producing less interference and allowing smaller, lighter batteries and/or longer periods of operation for portable equipment. The dissertation proposes a voice activity detection algorithm based on soft decisions concerning four parameters--energy level, spectral distribution, periodicity, and stationarity. The input signal to the voice activity detector is subdivided into 15 millisecond frames and each frame is processed and measured for relative energy content, percentage of signal energy at frequencies below 1000 Hz, and periodicity. Stationarity between and among frames is measured as a continuous variable using spectral covariance. The 15 msec frame is then classified as one of three types (voiced, unvoiced, or silence) based on a weighted sum of a series of probabilistic distance calculations performed on the values of the four parameters. The probabilistic distance calculations are nonlinear functions based on the distance between the measured values of a parameter and the mean values of the parameter for known voiced, unvoiced, and silence frames. The mean values are constantly updated during the speech stream and the weights are adaptive. A small series of comfort noise parameters are also added at the end of the speech stream to allow the receiver to replicate the background noise during the silent periods. The dissertation concludes with subjective and objective test results showing the performance of the proposed voice activity detection algorithm in various portable and mobile environments. The proposed algorithm is practical for mobile and portable wireless systems because it is simple enough to be implemented using a portion of a single DSP.