An Image Semantic Communication System Based on Swin Transformer and Generative Adversarial Networks
Wenyin Liang, Nan Ma, Linzi Shen, Mengshu Song, Xin Qi, Yiming Liu · 2024
To address the decline in image restoration quality at higher resolutions in semantic communication systems that utilize Convolutional Neural Networks (CNNs) as their backbone, researchers have increasingly focused on systems that incorporate the Swin Transformer. This model is recognized for its exceptional performance in the visual domain. However, most current semantic communication systems based on the Swin Transformer remain analog and do not quantize the encoder’s output, leading to incompatibility with digital communication systems. Therefore, this paper proposes a Swin Transformer and Generative Adversarial Network (GAN) based semantic system for image transmission named STGBS. By employing the Swin Transformer as the backbone network instead of the traditional CNN architecture, the system is better equipped to capture and extract both global and local features from images. Additionally, the incorporation of a quantizer ensures compatibility with digital communication systems, while adversarial training is utilized to enhance system stability and improve image reconstruction quality. Experimental results demonstrate that, compared to traditional methods and CNN-based Joint Source-Channel Coding (JSCC) systems, the proposed system shows significant improvements in Multi-Scale Structural Similarity Index (MS-SSIM), Peak Signal-to-Noise Ratio (PSNR), and overall image reconstruction quality.