End-to-end semantic preservation in text-aware image compression systems

Stefano Della Fiore, Alessandro Gnutti, Marco Dalai, Pierangelo Migliorati, Riccardo Leonardi · Signal Processing Image Communication · 2026

Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. Building on these insights, we further explore general-purpose encoders under extreme compression, investigating whether compact, visually degraded representations can still retain recoverable semantic information. Results demonstrate that semantic content can persist despite severe compression, bridging task-oriented text compression and broader machine-centered image coding.

Read the paper · More papers on PaperTik