OCR error detection and correction of an inflectional Indian language script
B.B. Chaudhuri, Umapada Pal · 1996
This paper deals with an OCR error detection and correction technique for a highly inflectional language script like Bangla (a major Indian language). This is the first report of its kind. Using two separate lexicons of root words and suffixes, candidate root-suffix pairs of each input word are detected, their grammatical agreement are tested and the root/suffix part in which the error has occurred is noted. The correction is made on the corresponding error part of the input string by a fast dictionary access technique. To do so some alternative strings are generated for an erroneous word. Among the alternative strings, those satisfying grammatical agreement in root-suffix and also having smallest Levenstein-Damerau distance are finally chosen as the correct ones. The system has an accuracy of 75.61%.