Multimodal Semantic Communication for 6G and Beyond: AI-Driven Architectures, Trends, and Challenges

Rowshan Mannan Oni, Fairuz Khan, Khalid Hasan Ador, Shahriar Hasan, Khan Mohammad Habibullah · IEEE Open Journal of the Communications Society · 2026

Since the inception of Shannon’s information theory in 1948, digital communication systems have primarily focused on reliable bit-level transmission under channel constraints. However, emerging 6G applications, characterized by massive connectivity and task-driven intelligence, challenge the efficiency and scalability of purely bit-centric paradigms, motivating a shift toward semantic communication, where task-relevant meaning rather than raw bit sequences is conveyed. In emerging 6G and beyond scenarios, distributed multi-agent systems rely on the exchange of heterogeneous multimodal data streams, necessitating semantic extraction and cross-modal fusion to construct coherent, task-relevant representations. Nevertheless, the existing literature on multimodal semantic communication remains fragmented, lacking a consolidated and systematic treatment. To address this gap, this paper presents a comprehensive systematic literature review of multimodal semantic communication, synthesizing and critically analyzing the literature across architectures, enabling technologies, and evaluation methodologies in a structured manner. A generic end-to-end multimodal semantic communication pipeline is derived from the reviewed literature, and the employed AI and ML architectures are analyzed across its constituent stages, covering modality-specific feature extraction, multimodal alignment, semantic fusion, and joint source-channel coding. In addition, knowledge base designs are categorized and assessed in terms of semantic fidelity, transmission overhead, and scalability. Furthermore, the channel models employed in the primary studies are critically examined, along with the applications of semantic communication in emerging 6G technologies and multi-agent systems. The synthesis of reported results indicates strong reconstruction performance of deep learning models even under low SNR conditions, and the adoption of generative models and large language models demonstrates notable gains in semantic fidelity and task effectiveness. Despite this progress, the field remains in its early stages, with significant open challenges in explainability, AI uncertainty, security, coexistence with conventional systems, and standardization.

Read the paper · More papers on PaperTik