A vocoder based on speech recognition and synthesis

Kechu Yi, Jun Cheng, An‐Liang Wang, Pu Zhang, Feng Liu, Weiying Li, Bin Yang, Shuanyi Du, Jun Gong · 2002

This paper introduces a speech recognition and synthesis based (SRSB) vocoder made by the authors, which has been judged by experts recently. The SRSB vocoder can encode Chinese speech of unlimited vocabulary at a bit rate of lower than 200 bps and reproduce speech with intelligibility of 95.2%. The vocoder consists of a real-time syllable recognizer to encode syllables based on composite hidden Markov modeling and a speech synthesizer to reproduce speech with syllable concatenation based on pitch-synchronous overlap-adding algorithm (PSOLA). It is capable of good prosodic modifications, since it can make use of prosodic parameters extracted from the input voice. Either terminal of it is implemented with an IBM-PC microcomputer equipped with a TMS320C30 DSP subsystem.

Read the paper · More papers on PaperTik