Text-Level Contrastive Learning for Scene Text Recognition
Junbin Zhuang, Yixuan Ren, Xia Li, Zhanpeng Liang · 2022
Scene text recognition (STR) is a task of identifying text from natural scene text images. Recently, based on the advantages of self-supervised contrastive learning, some studies incorporate contrastive learning strategies for STR task. However, these studies mainly focus on data argumentation of images from a visual perspective, ignoring the fact that scene text often contains large noise and has the characteristics of diversity. For addressing this issue, in this paper, we propose a text-level contrastive learning strategy to learn a better representation of text in scene text images to effectively improve the prediction performance of STR task. We perform extensive experiments on several public benchmark datasets and compare with baseline models, and the experimental results demonstrate the effectiveness of our method.