Screen-rendered text images recognition using a deep residual network based segmentation-free method
Xin Xu, Jun Zhou, Hong Zhang · 2018
Text images recognition has long been known as a research hotspot of computer vision. However, screen-rendered text image pose great challenges to current character or text recognition methods due to its low resolution and low signal-to-noise ratio properties. In this paper, a segmentation-free method utilizing Residual Network (ResNet) and Recurrent Neural Network (RNN)-Connectionist Temporal Classification (CTC) is proposed to recognize Chinese and English texts in screen-rendered images. Text lines are firstly extracted from screen-rendered images to obtain feature sequences. Then, a bidirectional RNN layer is applied to model the contextual information within feature sequences and predict identification results. Finally, a CTC method is employed to calculate loss and yield the results. The proposed method can achieve the best performance on ORAND-CAR-A dataset, ORAND-CAR-B dataset and a generated dataset with the recognition accuracy of 91.89%, 93.79% and 95.67%, respectively. Moreover, experiments on several real screen-rendered text images also demonstrate the effectiveness of the proposed method.