A CNN Postprocessor to Enhance Coded Speech
Ziyue Zhao, Samy Elshamy, Huijun Liu, Tim Fingscheidt · 2018
Postprocessors can be advantageously used to enhance trans-coded speech after the decoder on the receiver side. In this paper we present a convolutional neural network (CNN)-based postprocessor applying cepstral domain features to enhance the transcoded speech for various narrowband and wideband codecs without any modification of these codecs. Simulations show that the proposed postprocessor is able to improve the coded speech quality (PESQ or WB-PESQ) by 0.25 MOS-LQO points for G.711, 0.26 points for G.726, 0.81 points for G.722, and 0.2 points for AMR-WB. Moreover, a superior performance is observed for the proposed postprocessor compared to an ITU-T-standardized postfilter for G.711. We also show that AMR-WB at lower bitrates together with our proposed postprocessor is able to exceed the speech quality of AMR-WB at higher bitrates without postprocessing (up to 3 modes higher).