Improving text recognition by combining visual and linguistic features of text

Cong Tran, Khanh Nguyen-Trong, Cuong Pham, Dat Tran-Anh, Tien Nguyen-Thi-Tan · 2022

While being studied for several decades, Optical Character Recognition (OCR) has still been attracting considerable attention from researchers. Previous studies tend to focus on visual features of optical texts, such as texture, shape, and color to build OCR models. However, linguistic features, an important factor for OCR, has not been extensively investigated, especially for Vietnamese-OCR scanned documents. Therefore, we introduce a method to improve the performance of Vietnamese OCR by combining both visual and linguistic features of the optical text. The proposed method consists of (i) a domain-specific dictionary and (ii) a modified natural language processing model termed ABCNet, employed at the training and inference step, to determine the best candidate for the visual appearance of the text. Moreover, our method can easily be integrated with existing OCR methods to further increase their performance. Experimental results on a newly collected dataset show that the proposed method achieves an accuracy of 83.61% and a F1 score of 84.1%.

Read the paper · More papers on PaperTik