Lateral Cross-sectional Analysis based Classification of Bi-lingual Malayalam-English OCR using Distinct Features and non- Uniform Sub-classes
Bindu Philip, R. D. Sudhaker Samuel · 2008
extraction, to reduce dimensionality is a crucial process in the development of any robust optical character recognition system. In India bilingual documentation is very common especially government forms and formats, technical documents, postal documents, railways reservation forms etc. are bilingual and at times trilingual. Even documents printed in a single language often contain English words and numerals. An integrated recognition system that recognizes characters as well as numerals belonging to different languages in a single document has innumerous emerging applications and hence the motivation to work further in this area. Indian languages especially South Indian languages have several distinct characteristics which could be exploited to define features. In this paper Malayalam language is chosen as its script has exceedingly rich features. The length of the characters and the occurrence of transverse strokes in lateral directions turn out to be promising distinct features of these characters. This paper presents a lateral cross sectional analysis based approach for the recognition of Bi-lingual characters. Analysis is performed keeping track of the number of impulses at edges along each row in the image matrix resulting in distinct features. These features are grouped into subclasses based on heuristics depending on intra-subclass distances thus creating sub- classes of varying lengths. The feature selection method reduces computational complexity without compromising efficiency. A simple L2 - distance based classifier gives good performance.