A Swin Transformer-Based Approach to Semantic Transfer of Images

Bingkun Gan, Zhongdong Wu, Jingcong Gou, Pengbo Wang, Shangsi Ding · 2025

In recent years, visual models built on the Transformer architecture, such as the Swin Transformer, have achieved remarkable success in various image processing tasks. However, its core component, the multilayer perceptron (MLP) module, is prone to ignore local details when processing complex scenes, resulting in high-frequency information loss and insufficient multi-scale semantic synergy. To this end, this paper proposes an improved hybrid Attention-Enhanced MLP (AE-MLP) module, which leads to a new architecture called STA-JSCC. Finally, evaluation of the proposed method is conducted by integrating conventional metrics with feature similarity metrics. The obtained experimental outcomes demonstrate that the proposed STA-JSCC exhibits superior transmission effectiveness and communication content quality compared to other schemes, and thus the STA-JSCC approach has significant potential and advantages in enhancing the image reconstruction quality and semantic extraction capability.

Read the paper · More papers on PaperTik