Adaptive Dynamic Projection: A Novel Approach for Enhancing Multimodal Large Language Models
Pengfei Du, Chang Xu, Qian Zhang, Jian Zhai Wu, Kaiqi Zhao · Procedia Computer Science · 2025
Multimodal large language models (MLLMs) rely on projection mechanisms to bridge vision encoders and language models. Traditional projection techniques, including MLP-based and Resampler-based methods, often employ static token projections, resulting in inefficiencies during multimodal fusion. To address this challenge, we introduce Adaptive Dynamic Projection (ADP), a novel mechanism that adaptively selects and processes the most salient tokens, enhancing both computational efficiency and representational fidelity. ADP comprises three core components: Gated Token Selection, which identifies the most informative tokens, a Hierarchical Projection Network, ensuring streamlined token transformation, and an Adaptive Fusion Module, seamlessly integrating the dynamically selected visual tokens with textual representations. Through extensive evaluations on benchmark datasets such as ScienceQA and MMBench, ADP achieves state-of-the-art performance, attaining 94.61% accuracy on ScienceQA (surpassing the previous SoTA by 2.08%) and 79.1% on MMBench. Our method advances multimodal data integration through efficient token processing, delivering significant improvements in representational quality, with broad applicability across diverse vision-language understanding tasks.