Harmonic Model Based Excitation Enhancement for Low-Bit-Rate Speech Coding
Hong Kook Kim, Min Suk Lee, Chul Hong Kwon · IEICE Transactions on Information and Systems · 2004
SUMMARY A new excitation enhancement technique based on a har-monic model is proposed in this paper to improve the speech quality oflow-bit-rate speech coders. This technique is employed only in the de-coding process of speech coders and improves high-frequency componentsof excitation. We develop the procedure of harmonic model parametersestimation and harmonic generation and apply the technique to a currentstate-of-art low bit rate speech coder. Experiments on spectrum readingand spectrum distortion measurement show that the proposed excitationenhancement technique improves speech quality. key words: low-bit-rate speech coding, excitation enhancement, postpro-cessing, harmonic model based enhancement 1. Introduction In low-bit-rate speech coding, it is required to further en-hance spectral envelope and excitation for the improvementof speech quality due to the lack of assigned bits to both ofthem. Asoneofthesolutionsforspectralenvelopeenhance-ment, adaptive short-term postfilters constructed from thesynthesis filter [1] have been widely used, which could re-duce the audible distortion by enhancing the spectral peakswhile deemphasizing the spectral valleys. However, forCELP-type speech coders operating below 8kbit/s, enhanc-ing decoded excitation is crucial to decoded speech qual-ity because the decoded excitation is distributed sparsely sothat it often results in harsh quality speech sound. This phe-nomenon can be mitigated by applying a pulse dispersionfilter in the analysis stage[2] or in the synthesis stage[3].Nevertheless, we might still hear distortion in reconstructedspeech because the harmonic structure of the reconstructedspeechisnotaswellorganizedasitisintheoriginalspeech.A harmonic postfilteringapproach [4]anda long-term post-filter[1] were introduced to improve the harmonic structureof the reconstructed speech. However, these approachestendtomodifytheharmonicstructureofallfrequencybandsand thus the harmonic structure in low-frequency compo-nents of the reconstructed speech, which are usually wellmarked in speech coders, could be hurt after postprocess-ing.