Time-scale modification for speech coding
Dževdet Burazerovic, A.J. Gerrits, Rakesh Taori, Jean H. F. Ritzerfeld · TU/e Research Portal · 2000
Time-scale modification (TSM) of speech refers to compressing or expanding the timescale of speech, while preserving the identity of the speaker (pitch, formant structure). It can be used for various applications, such as text-to-speech synthesis, film/sound-track post-synchronisation, foreign language learning, etc. In this study, we have investigated ifTSM can be beneficial to speech coding. The central idea is to compress the time-scale of a speech signal prior to coding, enabling usage of less bits, and to expand it by a reciprocal factor after decoding, to come to the original time-scale. For this purpose, we have built a speech coding system, incorporating a compressor, an expander, and standard speech coders. As opposed to traditional approaches, we propose the use of parametric modeling of unvoiced sounds. For evaluation purposes, we have used 25 compression 33 expansion. Comparison with some standard coders is presented. We conclude that speech coders with TSM perform worse than dedicated speech coders at comparable bit-rates. However, TSM can be beneficial in providing graceful degradation at higher bit-rates.