Enhancing text extraction from natural scene images through diffusion models
Baojun Jia, Wenyu Zhang, Yifan Chen, Jin Ma · 2024
This paper focuses on super-resolution of natural scene text images, serving as a crucial preprocessing step for image recognition, with the aim of extracting text information from low-resolution images. In this preprocessing step, the primary challenges stem from the impact of factors such as lighting, occlusion, and distortion on text features in natural scenes, as well as the tendency of small pixel values in low-resolution images to be overlooked by large models. Previous works have primarily focused on the capability to extract features. However, the key to addressing this problem lies in restoring the different distributions between high-resolution and low-resolution images. As widely acknowledged, diffusion models excel in learning to generate distributions from datasets. Therefore, we introduce a conditional diffusion model to learn feature distributions among diverse images and apply histogram matching to address the issue of n patches. Our experimental results confirm the effectiveness of our approach.