Codec Network for Speech Bandwidth Extension

Chundong Xu, Xian-Peng Ling, Dongwen Ying · 2021

In order to further improve the performance of speech bandwidth extension, this paper proposes an end-to-end codec network model for speech bandwidth extension. It is modeled on the time-domain waveform and consists of encoder network, long short-term memory (LSTM) and decoder network. The encoder network is responsible for feature extraction and data dimensionality reduction of the input data, the LSTM is responsible for extracting the context-dependent information of the speech signal, and the decoder network is responsible for wideband speech reconstruction of the potential features output by the LSTM. In addition, this paper also proposes a time-frequency perception loss function to guide model training to generate more accurate time-domain waveforms and more realistic frequency-domain spectrum. Experimental results show that the reconstructed wideband speech generated by the model achieves better results in both subjective and objective evaluation.

Read the paper · More papers on PaperTik