A speech prediction model based on codec modeling and transformer decoding
Heming Wang, Yufeng Yang, DeLiang Wang · Computer Speech & Language · 2025
Speech prediction is essential for tasks like packet loss concealment and algorithmic delay compensation. This paper proposes a novel prediction algorithm that leverages a speech codec and transformer decoder to autoregressively predict missing frames. Unlike text-guided methods requiring auxiliary information, the proposed approach operates solely on speech for prediction. A comparative study is conducted to evaluate and compare the proposed and existing speech prediction methods on packet loss concealment (PLC) and frame-wise speech prediction tasks. Comprehensive experiments demonstrate that the proposed model achieves superior prediction results, which are substantially better than other state-of-the-art baselines, including on a recent PLC challenge. We also systematically examine factors influencing prediction performance, including context window lengths, prediction lengths, and training and inference strategies. • We propose a codec-based speech prediction approach that effectively leverages acoustic tokens and embeddings extracted from speech codecs. • We systematically evaluate the speech prediction performance of the proposed approach, and demonstrate that it outperforms recent baselines. • We investigate the factors that impact speech prediction performance and examine different training and inference strategies.