VLM-DM: Visual Language Models for Multitask Domain Adaptation in Driver Monitoring
Haozhuang Chi, Haohan Yang, Lie Yang, Chen Lv · 2025
Driver monitoring systems face critical challenges in modern transportation, including limited multitasking capabilities and a lack of interpretability. These limitations hinder the accurate and comprehensive assessment of driver states such as distraction, drowsiness, and emotions, which are essential to ensure road safety. This paper introduces visual language models for multitask domain adaptation in driver monitoring (VLM-DM), a novel framework that addresses these challenges by leveraging advanced visual language models for the simultaneous execution of multiple driver monitoring tasks. By employing parameter-efficient training methods such as Low-Rank Adaptation (LoRA) and integrating dynamic prompt tuning, VLM-DM achieves superior performance compared to state-of-the-art methods. Our experiments on three benchmark datasets across different driver states, demonstrating significant improvements in multitask accuracy and interpretability. This work highlights the potential of advanced multitask and multimodal architectures in developing robust, scalable, and interpretable driver monitoring systems for real-world applications.