Two-stage Reinforcement Learning Empowered Wireless Semantic Video Transmission

Mujian Zeng, Wenbo Ma, Kaipeng Zheng, Wenqin Zhuang, Mingkai Chen · 2025

With the rapid development of video conferencing technology, interactive multimedia services proliferate, resulting in a surge in business traffic. Remote work and online collaboration will become the mainstream office in the future, which poses new challenges to computer vision technology in the field of communication. At the same time, the appearance of semantic communication has solved the problem that the traditional video coding technology can not deal with semantic redundancy. Therefore, we propose the Semantic Transmission Optimization of Video Conference(STOVC) architecture to meet the interactivity requirements of multimedia services. We first use Iterated Integrated Attribution(IIA) and Segment Anything Model 2(SAM2) to understand the video and complete the semantic segmentation of the video. We then propose Neural Video Compression with Actor-Critic(NVCAC) to optimize video encoding and transmission strategies to balance video compression with semantic fidelity. Finally, the$\alpha$-fusion method is used to reconstruct the video at the receiving end. In addition, we use DeepCache accelerated stable diffusion to generate conference video to enrich the data set. The experimental results show that STOVC has a better performance than H.264 and H.265 at low bit rate.

Read the paper · More papers on PaperTik