Text-and-timbre-based speech semantic coding for ultra-low-bitrate communications

Yang Xiaoniu, Liping Qian, Lyu Sikai, Qian Wang, Wang Wei · China Communications · 2025

To address the contradiction between the explosive growth of wireless data and the limited spectrum resources, semantic communication has been emerging as a promising communication paradigm. In this paper, we thus design a speech semantic coded communication system, referred to as Deep-STS (i.e., Deep-learning based Speech To Speech), for the low-bandwidth speech communication. Specifically, we first deeply compress the speech data through extracting the textual information from the speech based on the conformer encoder and connectionist temporal classification decoder at the transmitter side of Deep-STS system. In order to facilitate the final speech timbre recovery, we also extract the short-term timbre feature of speech signals only for the starting 2s duration by the long short-term memory network. Then, the Reed-Solomon coding and hybrid automatic repeat request protocol are applied to improve the reliability of transmitting the extracted text and timbre feature over the wireless channel. Third, we reconstruct the speech signal by the mel spectrogram prediction network and vocoder, when the extracted text is received along with the timbre feature at the receiver of Deep-STS system. Finally, we develop the demo system based on the USRP and GNU radio for the performance evaluation of Deep-STS. Numerical results show that the accuracy of text extraction approaches 95%, and the mel cepstral distortion between the recovered speech signal and the original one in the spectrum domain is less than 10. Furthermore, the experimental results show that the proposed Deep-STS system can reduce the total delay of speech communication by 85% on average compared to the G.723 coding at the transmission rate of 5.4 kbps. More importantly, the coding rate of the proposed Deep-STS system is extremely low, only 0.2 kbps for continuous speech communication. It is worth noting that the Deep-STS with lower coding rate can support the low-zero-power speech communication, unveiling a new era in ultra-efficient coded communications.

Read the paper · More papers on PaperTik