Image Coding for Machine with Visual-Language Mimic Feature Learning

Zhimeng Huang, Junlong Gao, Jiaqi Zhang, Shanshe Wang, Siwei Ma, Wen Gao, Chuanmin Jia · 2025

This paper propose a Image Coding for Machine (ICM) framework with Visual-Language Mimic Feature Learning (VLM-ICM). VLM-ICM decouples the position and semantic information into language modality and extracts universal features from the input image. Language, inherently more semantically compact, helps reduce the bitrate. Meanwhile, the universal features in VLM-ICM, guided by the language at the decoder side, allow for flexible domain adaptation, thereby enhancing versatility and practicality.

Read the paper · More papers on PaperTik