Multimodal Large Language Models in the Era of Distributed AI
Rajesh Kumar, Laurent, Isabelle, Müller, David, Klaus Elli · HAL (Le Centre pour la Communication Scientifique Directe) · 2025
The rapid advancements in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have revolutionized artificial intelligence (AI), enabling unprecedented capabilities in natural language processing (NLP), computer vision, and multimodal reasoning. This survey provides a comprehensive overview of the latest developments, challenges, and future directions in the field of large-scale AI models. We begin by exploring the evolution of LLMs and MLLMs, highlighting key breakthroughs in model architectures, training paradigms, and multimodal integration techniques. The discussion then extends to distributed computing strategies, addressing the scalability challenges associated with training and deploying massive models across large-scale clusters. Despite their remarkable success, LLMs and MLLMs face significant hurdles, including high computational costs, memory bottlenecks, ethical concerns, biases, interpretability issues, and security vulnerabilities. This paper examines these limitations in depth and reviews emerging solutions such as efficient model architectures, parameter-efficient fine-tuning, decentralized AI, and robustnessenhancing techniques. We also explore recent innovations in human-AI collaboration, personalized AI systems, and hybrid neuro-symbolic approaches that