OCR for Malayalam script using neural networks
M. Abdul Rahiman, M. S. Rajasree · 2009
This paper specifies an OCR system for printed Malayalam characters. Malayalam is the principal language of the South Indian state Kerala. The input to the system would be the scanned image of a page of text and the output is a machine editable file. Malayalam Character recognition is a complex task because of the presence of two scripts; old script and new script and a lot of combinational characters. Initially, the image is preprocessed to remove noise. Then skew correction methods are applied to the document. Lines, words and characters are segmented from the processed document image. The proposed method uses wavelet analysis for extracting features of the image and Back propagation neural network is used to accomplish the recognition tasks.