Statistical treatment of the acoustic effects of vocal effort

Kung-Pu Li, Peter Benson · The Journal of the Acoustical Society of America · 1985

Speech recognition performs best when trained and operated in the same environment. Such training is often impractical, especially in military applications. Apart from the difference in the noise environment, talkers change their speech habits when speaking over poor quality channels. The acoustic changes due to increased vocal effort reduce the performance of speech recognizers. While the acoustic effects of vocal effort are varied, it is possible to deal with a bulk of the differences using a statistical approach. A database of two talkers was collected under conditions which produced a variety of degrees of vocal effort. These conditions included both loud sidetone and noise which, when delivered over headphones, change stress and vocal effort. The talkers differed in their reaction to the extreme level in that one shouted and the other did not. Read materials consisted of lists of digits, randomized but the same across talkers. Acoustical parameters such as pitch, amplitude, and word duration were measured and compared across conditions. Using standard statistical procedures, transforms were computed to map the speech from several conditions into a common transformed space. Recognition experiments demonstrated that significant increases in automatic speech recognition performance can be achieved when templates from one condition are matched against unknown speech, when both have been transformed into the common space. An analysis of the recognition errors was made in light of the acoustic differences between vocal effort conditions.

Read the paper · More papers on PaperTik