Handwritten and machine printed text separation from Kannada document images
Rajmohan Pardeshi, Mallikarjun Hangarge, Srikanth Doddamani, KC Santosh · 2016
Handwritten and machine printed (H&P) text separation from document images is a precursor to advance the performance of the OCR system. This paper demonstrates the competence of frequency domain features for the classification of H&P text words. We propose wavelet-like discrete cosine transform (WDCT) based features. We conduct an experiment on a large dataset of 2000 text words of popular south Indian script Kannada, where k-NN classifier is employed. The efficacy of frequency domain features is experimentally validated with the classification accuracy of WDCT 99.50% using ten fold cross validation.