Multi-User Semantic Communication for Interactive Speech Dialogue With Dynamic Resource Allocation
Zhikai Liu, Tharmalingam Ratnarajah · IEEE Transactions on Network Science and Engineering · 2025
This paper presents a novel semantic communication model for multi-turn interactive speech dialogue tasks, enabling multiple users to engage in dialogue with a language model deployed at a base station (BS) in a massive MIMO (mMIMO) system. The framework utilizes a time-slotted design and integrates a distributed memory module to store semantic speech features from previous interactions. The speech quantization model incorporates a token fusion mechanism to address token length and computational complexity. An enhanced speech context retrieval model, augmented with knowledge distillation, is introduced to retrieve relevant historical contexts from the memory module efficiently, significantly improving response quality. Long short-term memory (LSTM) network-based joint source-channel (JSC) encoders and decoders are employed to mitigate multi-user interference. The system undergoes a comprehensive four-step training process, optimizing each component sequentially. In the final phase, Randomized Ensembled Double$Q$-Learning (REDQ), a reinforcement learning algorithm, is applied to optimize bandwidth and power allocation, effectively reducing latency and enhancing the accuracy of generated responses. Extensive simulations validate that the proposed semantic communication system significantly improves answer accuracy, particularly in multi-user scenarios, while effectively reducing computational complexity and latency.