Fuse Before Transmit: A Multimodal Semantic Communication Method
Leiyu Wang, Xiaodong Xu, Hao Chen, Yaping Sun, Nan Ma, Ping Zhang · 2024
Semantic communication has been recognized as a key enabling technology for future 6G wireless networks. Existing studies have predominantly focused on single modal semantic communication, such as the design of image and text encoding and decoding for transmission. However, under poor channel conditions, single modal semantic data is easily affected, leading to semantic distortion at the receiver end. To address this limitation, this paper proposes a multimodal semantic communication (MMSC) framework. At the transmitter, MMSC consists of an image encoder, a text encoder, and a multimodal semantic feature fusion unit. The multimodal semantic feature fusion unit integrates these modalities into a unified semantic feature vector for transmission. At the receiver, the transmitted multimodal feature vector can be simultaneously sent to the image decoder and the text decoder for recovery. Experimental results demonstrate that the proposed MMSC approach can achieve efficient reconstruction of multimodal data.