A very low bit rate speech coder using HMM with speaker adaptation

Takashi Masuko, Keiichi Tokuda, Takao Kobayashi · 1998

This paper describes a speaker adaptation technique for a phonetic vocoder based on HMM. In the vocoder, the encoder performs phoneme recognition and transmits phoneme indexes and state durations to the decoder, and the decoder synthesizes speech using HMM-based speech synthesis technique. One of the main problems of this vocoder is that the voice characteristics of synthetic speech depend on HMMs used in the decoder, and are therefore fixed regardless of a variety of input speakers. To overcome this problem, we adapt HMMs to input speech by transmitting transfer vectors, information on mismatch between the input speech and HMMs. The results of the subjective tests show that the performance of the proposed vocoder without quantization of transfer vectors is comparable to that of a speaker dependent vocoder. 1. INTRODUCTION To code speech at rates on the order of 100 bit/s, phonetic or segment vocoders are the most popular techniques [1]-[6]. These coders decompose speech into a seque...

Read the paper · More papers on PaperTik