Audio/speech coding using the matching pursuit with frame-based psychoacoustic optimized time-frequency dictionaries and its performance evaluation
Alexey Petrovsky, Vadzim Herasimovich, Alexander Alexandrovich Petrovsky · 2016
This paper presents an audio/speech coding algorithm using the matching pursuit with the dynamic dictionary forming based on wavelet packet decomposition and its performance evaluation. The proposed methodology for selecting the most relevant wavelet coefficients is based on maximizing the matching between the auditory excitation scalograms associated with the original and the modeled signal correspondingly. The major advantage of this method is that the wavelet packet dictionary is perceptually optimized for each signal segment. It is obtained to reduce the number of the coefficients required to achieve a given perceptual distortion. Flexibility of the presented algorithm gives a possibility to change bitrate depending on the transmission channel limitation. Objective evaluation of the reconstructed audio signals by PEMO-Q model and comparison with the modern popular audio encoders such as Opus and Vorbis is provided. Received results show the high quality of the reconstructed signal of the proposed audio/speech coder.