Speech Coding Using an Enhanced Sinusoidal Model at Low Bit-Rate
Yasheng Qian, Jia Liu, Feng Chongxi, Peter Kaba · 1992
An enhanced sinusoidal model, which employs the time-varying amplitudes of three components to track the fast dynamicd variations during the transition speech segments. and exploits the redundancies between the near-neighborhood components to reduce the number of sinusoidal components to a maximum of ?O with high synthesized quality is presented. Many components can be determined by linear prediction of the dominant and fundamental components, thereby reducing the number of the parameters required to be transmitted and the corresponding bit rate. This approach improves the synthesized quality of the unvoiced and transition speech segmenrs. An optimal algorithm for extracting dominant frequencies by formats and pitches is compared with a DFT method. The effects on the synthesis quality of the number of the time-varying amplitudes and the different base functions are compared. Two vector quantization codebooks with group classifications are developed to reduce the storage and computation load for a 4.8 kbitsls coder. Objective meuurernents give a cepstrurn distance of 2.62 dB for several phonetically balanced sentences. Informal listening tests have shown that the proposed speech coder with an enhanced sinusoidal model can obtain good quality speech at 4.8 kbits/s.