A phonetic vocoder with scalable adaptation to speaker codebooks
Israel Halaly, Y. Bistritz · 2008
The paper presents a very low bit rate phonetic vocoder based on speech recognition and synthesis with scalable speaker adaptation using a set of speaker phoneme codebooks (SPCBs). These SPCBs are designed by an LBG-like algorithm and are available to both encoder and decoder. The encoder compares periodically the input speech to adapted synthesized speech produced by the SPCBs. For each SPCB, the sequence of recognized phonemes and the pitch contour is used to synthesize speech that is adapted by a scalable multi-segmented spectral warping. The information about the SPCB and its adaptation parameters that achieve the least spectral distortion is transmitted to the decoder. This revises a vocoder presented earlier this year by considering for it a more flexible adaptation scheme with just a slight increase of the bit rate. Experiments held at a typical low bit rate of phonetic vocoders (around 300 bps) demonstrate that the scalable adaptation reduces the average spectral distortion and improves speaker recognizability as judged by listeners.