A Differential Quantization Based End-to-End Neural Speech Codec

Pingli Lu, Liang Xu, Jing Wang · 2024

Speech codecs efficiently compress speech signals, reducing the bandwidth occupied during communication. With the development of neural networks and deep learning, end-to-end speech codecs based on neural network structures have emerged. Compared to traditional codecs, these neural speech codecs can reconstruct higher-quality speech at lower bitrates. However, the performance of neural speech codecs drastically deteriorates when the communication bitrate drops to 1 kbps or below, as these codecs are based on residual quantization, which has limited performance at low bitrates. In this paper, a differential quantization based neural speech codec is proposed. In particular, the quantization focuses on the importance of difference frames and preserves key information with as few bits as possible. Meanwhile, we propose a compensator to further improve the reconstructed speech quality. Both subjective and objective evaluations demonstrate that our proposed method can achieve a higher quality of reconstructed speech at 0.6 kbps than Sound-Stream at 3 kbps. The entire model is causal, supporting streaming and real-time inference.

Read the paper · More papers on PaperTik