Neighborhood-aware adapter for vision-language models

Luchen L. Ji, Xianhui Wang, Xiaolin Xu, Peiyu Lu, Xiaoxu Li · 2025

Pre-trained vision-language models (VLMs) have demonstrated outstanding performance in various downstream tasks, but conventional adapter methods, relying on fully connected layers, struggle to capture long-range dependencies and local features. To overcome this limitation, we propose the Neighborhood-Aware Adapter (NAA), which enhances VLM generalization by emphasizing the relationships between neighboring feature regions during fine-tuning. NAA improves feature embeddings by reinforcing neighborhood correlations, enabling the model to learn complex spatial patterns. Our experiments show that NAA outperforms existing methods, achieving a 4.43% accuracy improvement on ImageNet and significantly boosting zero-shot learning and cross-dataset generalization. Ablation studies validate the contributions of each NAA component.

Read the paper · More papers on PaperTik