Foundation Model-Aided Channel-Adaptive Video Semantic Communication and Prototype Validation

Jiarun Ding, Peiwen Jiang, Chao-Kai Wen, Xiao Li, Shi Hong Jin · IEEE Transactions on Wireless Communications · 2025

The increasing demand for services such as live streaming and virtual reality places significant pressure on wireless communication systems. Enhancing system performance or reducing bandwidth consumption is critical for delivering high-quality video experiences. Semantic communication, which focuses on the transmission of meaning, offers a promising solution. However, existing approaches are often limited to single scenarios, rely on simple channels, lack adaptability to dynamic wireless environments, and remain untested in practical air interfaces. To address these challenges, we propose a foundation model-aided universal video semantic communication framework designed for pixel-wise reconstruction across diverse scenarios. This framework enables the transmission of entire videos using joint source-channel coding (JSCC) based on optical flow estimation and leverages multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) for efficient semantic delivery in 3rd generation partnership project (3GPP) standard channels. In scenarios requiring full transmission for regions of interest and selective transmission for other areas, the framework employs a foundation model for segmentation, followed by JSCC and delivery. Furthermore, we introduce a channel condition number-adaptive semantic remapping method based on an attention mechanism to mitigate the effects of wireless fading. To validate our approach, we implement the framework on a testbed and develop two online demonstrations. Simulations and over-the-air experiments confirm significant improvements in video quality and substantial reductions in bandwidth overhead compared to existing methods.

Read the paper · More papers on PaperTik