Enhancing Vision-Language Models Through Passive Online Target Recognition and Test-Time Adaptation

Lijuan Chen, M. Yu, Hao Yang Shen, J. S. Luo · 2024

In the realms of computer vision and natural language processing, researchers have aimed to effectively integrate images and text. Traditional methods, which depend on feature extraction and classifiers for target category inference, often fail to model the association between images and textual descriptions, resulting in insufficient information transfer when treating image processing and semantic inference separately. A novel approach based on vision-language large models offers an end-to-end learning framework that simultaneously processes and integrates image and textual information. This method enhances target recognition accuracy and robustness, exhibits strong zero-shot capabilities, and generalizes well across various image types. Furthermore, passive online target recognition, a form of unsupervised domain adaptation, eliminates the need for additional labeled data or external inputs by relying on target domain data for adaptation and learning. This technique automatically learns the feature distribution and structure of the target domain, enabling effective knowledge transfer from the source domain model to the target domain, even without source domain label data.

Read the paper · More papers on PaperTik