Yi Characters Recognition Based on Tesseract-OCR
Peiyu Sun, Qiuyan Xie, Zhaokang Wu, Xiaoyu Feng, Jiajun Cai, Yulian Jiang · 2019
OCR(optical character recognition) is a general image recognition technology, which has been widely used in Tibetan and Mongolian minority characters recognition, but it is seldom used for Yi characters recognition. In order to promote the exchange of Yi and Chinese characters, this paper proposed the Yi characters recognition system based on OCR . This method uses the open source Tesseract-OCR engine. The system uses the jTeesBox Editor to extract the frame of the Yi characters, obtain the feature of the font, thus correct the characters, and then train the input Yi characters through Tesseract-OCR. In this paper, 900 pictures of Yi characters are used for training. The experimental results show that the recognition accuracy of the method reaches 85.9% , which has the good effect.