Characteristic behavior of long-term speech vector sums with application to speaker identification

Timothy R. Thomas, Wojciech J. Zakrzewski · The Journal of the Acoustical Society of America · 1988

It is known that the parameters derived from linear predictive coding of speech can provide a basis for speaker identification, and that long-term averages improve the reliability of such recognition schemes. The present study obtained long-term descriptors of speech by converting 20-ms frames into unit direction vectors in 14-dimensional cepstral coefficient space. Various numbers of these vectors were then summed, either sequentially or randomly, and the statistical characteristics of the lengths and directions were explored, using both analytical and computational tools. Interspeaker and intersession comparisons were then made, as well as comparisons to sums obtained from randomly generated vectors and from frames selected on the basis of total energy. Several interesting findings were revealed, including the observation that while the average angle between frames and the frame-to-frame changes in direction remained remarkably consistent across speakers and sessions, the direction of the long-term sums was reliably different between speakers, but not sessions. Furthermore, a proper selection on the basis of total energy in the frame can enhance interspeaker differences.

Read the paper · More papers on PaperTik