From text to multimodality: A comprehensive survey of LLMs and MLLMs across diverse domains

Sampath Anbukkarasi, Hemalatha S, Arunkumar B, Shanmugavalli V, Kokila S, Yashaswini K A · ICT Express · 2026

The rapid development of Artificial Intelligence has made it possible for advanced language learning systems, including Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), to be created. These models go beyond the traditional text-based processing by incorporating different modalities like video, audio, text creating more accurate prediction systems. Here, we aim to give a thorough look at multimodal models focusing on their use and challenges in different real-world domains including healthcare, education, industry, governance, security, assistive systems, scientific research, and creative media. We break down model architecture, multimodal fusion techniques, training methods, and specific adaptations for domains that improve understanding, decision-making, and interactive capabilities. On top of that, we discuss major benchmark datasets, performance evaluation methods, and real-world examples with the help of which we point out the present capabilities and practical limitations. We engage with major challenges such as data inconsistency, model alignment scalability computational cost, privacy issues, bias, and explainability. Besides, this article explores instructional-tuned multimodal systems, lightweight domain adaptations, retrieval-augmented multimodal reasoning, and human-in-the-loop learning models that are gaining ground. We finish by listing prospective research directions that lead to trustworthy, energy-efficient, context-aware, and ethically aligned multimodal AI systems that connect artificial and augmented intelligence. This article is a thorough guide for researchers, professionals, and policymakers in pushing forward multimodal models to strong and socially responsible usage across different domains.

Read the paper · More papers on PaperTik