Driver-Guide: A Multimodal Large Language Model-Based Agent for Driving Scene Understanding
Yabing Ran, Bingzhao Gao, Qiankun Yu · 2025
The research proposes a driving scene understanding agent based on a Multimodal Large Language Model (MLLM) to enhance scene understanding and human-machine interaction capabilities in autonomous driving. A high-quality multimodal driving scene dataset has been constructed, covering tasks such as scene description, question answering, and action suggestions. Two types of multimodal feature mappers (MLP with two hidden layers and Q-Former) were compared, and experiments demonstrated the superior performance of Q-Former in complex scene. Additionally, by incorporating a traffic rules knowledge base and Retrieval-Augmented Generation (RAG), the system effectively reduces hallucination phenomena, ensuring the accuracy and compliance of model responses. Experimental results indicate that the proposed system excels in complex scene semantic understanding and interaction tasks, offering a novel solution to address the limitations of traditional methods in complex driving scene understanding.