A 0.75 kbps speech codec using recognition and synthesis schemes
Heng-Chou Chen, Chin-Yung Chen, Kai-Ming Tsou, O.T.-C. Chen · 2002
In this paper, we proposed a very low bit-rate speech codec using recognition and synthesis schemes. The 2512 speech units, including 48 phones and 2463 diphones are utilized in the recognition process. The three-state continuous hidden Markov model, excluding the start and final states, is applied to model these speech units. In addition to the recognized phonetic index, the corresponding phonetic frame length is also the compressed information. In order to obtain a better quality of the reconstructed speech, pitch periods and pitch gains are realized to preserve the speaker's personal characteristics. In the synthesis process, the time-domain pitch-synchronous overlapped addition scheme is utilized to synthesize a high-quality speech waveform. In our tests, a more than 90% recognition accuracy can be achieved when the user speaks in a normal behavior. The reconstructed speech quality can be above a mean opinion score of 3.0, and a diagnostic rhyme test score of 92.