An extended method for recognition of broken typewritten characters special reference to tamil script
Nirase Fathima Abubacker, Raman Indra Gandhi · 2011
Preparing clean and clear images for the recognition engines is often taken for granted as a trivial task that requires little attention. Most of the existing OCRs have been designed in such a way that which correctly identify fine printed documents in all scripts. The performance of standard machine printed OCR system works fails, if it is tested on documents with distorted characters. This paper presents an approach to overcome the difficulties presented in such distorted type written documents especially with broken characters. As a first step, isolation of character is forwarded using character position location and character localization and enclosing it in a matrix which will be analyzing and repairing in the later part of our study. An attempt is incorporated using shape and line tracing method for recognition of distorted broken characters and then it is fine tuned by lexical knowledge.