VAE-Based Joint Source Channel Coding with Residual Swin Transformer for Facial Image Communication
Ranhe Zhang, Hongbin Ma, Yingli Wang · 2025
This paper presents a novel joint source-channel coding (JSCC) framework for robust communication of facial image semantics over noisy wireless channels. The system uses a variational autoencoder (VAE) to extract low-dimensional latent representations that capture the essential semantic features of facial images. These latent vectors are then mapped to channel inputs using a multi-layer perceptron (MLP), eliminating the need for separate source and channel coding. On the decoder side, we introduce a residual Swin Transformer integrated with sampling-based convolutional layers to enhance reconstruction quality by capturing both local textures and global semantic structures. Experiments conducted on the FFHQ dataset over AWGN and Rayleigh fading channels demonstrate that our approach outperforms existing DeepJSCC and SemVit methods in terms of PSNR and perceptual quality, particularly under low SNR conditions. The proposed architecture effectively balances semantic fidelity and transmission robustness, making it suitable for future intelligent wireless communication systems.