Nonlinear Time Compression and Time Normalization of Speech
Suzanne L. Hanauer, Manfred R. Schroeder · The Journal of the Acoustical Society of America · 1966
In a study of the effects of different speaking rates on articulator dynamics [J. E. Miller, J. R. Pierce, and M. V. Mathews, J. Acoust. Soc. Am. 34, 1978(A) (1962)], it was found that, when talkers were instructed to speak faster, they shortened the “steady-state” vowel portions while keeping the phoneme transitions approximately constant in duration. Thus, in speaking faster, humans use nonlinear time compression. In the present nonlinear time-compression scheme, speech is separated into contiguous narrow frequency bands, for each of which signals representing amplitudes and instantaneous frequencies are derived [J. L. Flanagan et al., J. Acoust. Soc. Am. 38, 939(A) (1965)]. The sampling rate of these signals is made proportional to the total temporal derivative of the logarithmic short-time spectrum. As a result, the sampling rate is higher during phoneme transitions than during “steady-state” portions. Nonlinear time compression is achieved by applying these nonuniform time samples at a constant rate to a speech synthesizer. Thus, speech can be time-compressed in a manner resembling human speech at increased speaking rates. Results are compared with linear time compression for the same average compression factor. Also, the application of nonlinear time compression to time normalization for automatic speech recognition is discussed.