A Model of Vietnamese Optical Character Recognition

Kha Tu Huynh, Cong-Man Tran, Huu Sy Le · 2022 RIVF International Conference on Computing and Communication Technologies (RIVF) · 2022

Optical Character Recognition (OCR) is a method to transform images in words into digital documents in the computer vision field. This helps digital or hand-written words and characters in images to be recognized by reading a document file with a monitor. However, a computer can only understand a picture as pixels or a tree-dimension array with values from 0 to 255. OCR applications can help translate nearby pixels into characters, words, and sentences. In this paper, we propose a transformer model to solve the Vietnamese OCR problem which is optimized or shortened to fit on the GPU while still producing solid results and loaded pre-trained from the HuggingFace Hub [1]. The proposed model achieves the maximum level of accuracy, 96.2%, with a CER of 0.8% in case of training on a labeled dataset that could not discriminate between single and compound words. The simulations result also proves that the training and assessment losses are reduced quickly in the first half and steadily in the second half. Due to the complexity of the Vietnamese and a few studies related to identifying Vietnamese through images, our study can be considered as an effective and supportive model for optical character recognition and a basis for related research.

Read the paper · More papers on PaperTik