QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a Query-aware Transformer
Chongyu Liu, Qing Jiang, Dezhi Peng, Yuxin Kong, Jiaixin Zhang, Longfei Xiong, Jiwei Duan, Cheng Sun, Lianwen Jin · Neurocomputing · 2024
Scene text image super-resolution (STISR) has obtained widespread attention in recent years due to its ability to enhance text recognition performance. Many previous methods proposed to incorporate text prior knowledge into the super-resolution architecture for reconstructing high-quality text images. However, these text priors are typically derived from pretrained text recognition models, and the inaccurate recognition feedback will hinder overall performance. In this paper, we propose a novel model, QT-TextSR, which promotes scene text image super-resolution by introducing efficient interaction with text recognition to release the inaccurate text feedback through a Query-aware Transformer. Specifically, QT-TextSR decomposes scene text image super-resolution and scene text recognition into different sets of queries within a Vision-Language Cooperation Module, explicitly modeling discriminative and interactive features between text recognition and text image super-resolution tasks. By employing two separate yet simultaneous projection heads on the corresponding features, QT-TextSR can recover the low-quality text image meanwhile obtain the recognition results. Additionally, to mitigate the limitations caused by recognition errors and enhance text structure preservation, we introduce a strong texture prior through self-supervised pre-training, leveraging visual cues more effectively. Experiments on public dataset, TextZoom demonstrate that our QT-TextSR significantly outperforms previous state-of-the-art methods in the metrics of Recognition Accuracy (68% v s . 65.5%), PSNR (22.51 v s . 22.10), and SSIM (0.7960 v s . 7930). The code for QT-TextSR is available at https://github.com/lcy0604/QT-TextSR .