An Initial Study to Solve Imbalance Sundanese Handwritten Dataset in Character Recognition
Erick Paulus, Mira Suryani, S. Hadi, Fadhliyah Natsir · 2018 Third International Conference on Informatics and Computing (ICIC) · 2018
The ancient Sundanese manuscripts are one of the tangible cultural heritage in Indonesia that contain various local wisdoms, medicine formulae, and social life stories from the past. In order to preserve those intangible assets, the society need to retrieve the information written in the manuscripts. One way to retrieve this information is by developing specific OCR for handwriting text. The basic of OCR process needs to recognize the glyphs. Not only the recognition architecture model, but the quantity and quality of glyph images can also be real challenges for glyph recognition. Since our ancient Sundanese manuscript dataset has a few samples of each class character, we proposed an initial study to create automatic synthetic data generator to deal with it. The original of isolated Sundanese Character have been multiplied using several images. Then, synthetic datasets evaluated using KNN classification with histogram of oriented gradient as feature extraction. The results of experimental study show that our proposed synthetic data generator increase the recognition rate at roughly 77%.