A Method of Synthesizing Handwritten Chinese Images for Data Augmentation
Xi Shen, Ronaldo Messina · 2016
The performance of printed document recognition has been significantly improved by generating synthetic images to augment the training data, particularly by providing more variability in the linguistic contents. Handwriting recognition benefits less from this data augmentation and the only variability that is usually added is via artificially generated combinations of skew, slant and noise. Generating handwritten text is complex due to variations in form, scale and spatial placement of the characters, and can be further complicated by the cursive aspects of the script. We propose a novel strategy, in the particular case of Chinese characters, to generate synthetic lines of text, given samples of the isolated characters. The well-known CASIA database is used to train MDLSTM-RNN models and also in the creation of synthetic line images. On an independent set of real-world images, a model trained only on synthetic images achieved a small relative reduction of 4.4% in the character error rate with respect to a baseline model trained exclusively on real images, while training on a combination of real and synthetic images resulted in a appreciable reduction of 10.4%.