Video Token Sparsification for Efficient Multimodal LLMs in Driving Visual Question Answering

Yunsheng Ma, Amr Abdelraouf, Rohit Gupta, Ahmadreza Moradipari, Ziran Wang, Kyungtae Han · 2025

Multimodal large language models (MLLMs) have shown significant potential in enhancing driving scene understanding and visual question answering (VQA) through advanced logical reasoning capabilities. These tasks support driving action generation and explanation, especially in end-to-end autonomous driving applications. However, deploying these models poses a significant challenge due to their substantial parameter sizes and computational demands, which often exceed onboard computational limits. A key limitation stems from the large number of visual tokens needed to capture detailed, long-context visual information, resulting in increased latency and memory use. To address this, we propose Video Token Sparsification (VTS), a novel approach that leverages redundancy in consecutive video frames to reduce visual tokens while preserving critical information. VTS employs a lightweight CNN-based model to identify key frames and prune less informative tokens, mitigating hallucinations and boosting inference throughput without performance loss. Comprehensive experiments on the LingoQA and DRAMA benchmarks show that VTS achieves up to a 33% improvement in inference throughput and a 28% reduction in memory usage compared to baselines, maintaining comparable performance.

Read the paper · More papers on PaperTik