How Accurate Can Large Vision Language Model Perform for Images with Compression Degradation?

Xiaohan Fang, Peilin Chen, Meng Wang, Shiqi Wang · 2024

The rapid evolution of Large Language Models (LLMs) has spurred the development of Large Vision Language Models (LVLMs), which demonstrate remarkable proficiency in various computer vision tasks through corresponding input prompts. These models perform impressively on diverse multimodal benchmarks, matching the effectiveness of conventional task-specific models. Nonetheless, many experiments with LVLMs often assume that the images are pristine, overlooking potential information loss during image transmission. To explore how image compression affects the semantic analysis capabilities of LVLMs in real-world scenarios, this paper presents a new imagetext dataset named GPT-COMP. This dataset comprises 80,000 natural scene images, including 20,000 raw images from two public datasets. Each image is subjected to three different levels of compression distortion (QP = 32, 42, 52) using the latest Versatile Video Coding (VVC) Test Model. We leverage these variably compressed images in GPT-COMP to evaluate the state-of-the-art LVLM, GPT-4o, in vision understanding tasks. The text responses generated by GPT-4o are further organized into the GPT-COMP dataset and serve as the basis for evaluation. Specifically, the scene understanding capabilities regarding different compression levels are measured based on extracted semantic features and our proposed self-scoring strategy. This analysis sheds light on how image compression affects the semantic analysis capability of LLMs, offering valuable insights into the resilience of these models under realistic, suboptimal conditions.

Read the paper · More papers on PaperTik