LLM performance in multimodal learning environments: study of integration of text with visual, audio, and sensor data for holistic decision-making
Nikunj Agarwal, Aditi Choudhary, Aditya K. Gupta, Pulkit Jain, Mukund B. Wagh, Dinesh Besiahgari · 2025
The advent of Large Language Models (LLMs) has redefined the boundaries of artificial intelligence, particularly in natural language processing. With their remarkable ability to generate coherent text, LLMs are now being explored for their potential in multimodal learning environments where data from text, visual, audio, and sensor inputs converge. This study delves into the integration of these modalities using LLMs, focusing on their performance in holistic decision-making tasks. By analysing foundational principles, current methodologies, and future directions, this paper aims to provide a comprehensive understanding of the opportunities and challenges in leveraging LLMs for multimodal applications. Empirical insights are supported with technical details, evaluation metrics, and real-world applications.