Offline Recognition Of Image for content Based

Smeet D. Thakur, Smita S. Sikchi · 2013

This document gives formatting instructions for authors preparing papers for publication in the Proceedings of an IEEE conference. The authors must follow the instructions given in the document for the papers to be published. You can use this document as both an instruction set and as a template into which you can type your own text. Keywords — Pre-processing, 2. Segmentation. 3. Feature Extraction, 4.Recognition, 5. Classification I. INTRODUCTION OCR work on printed Devnagari script started in early 1970s. Earlier studies on Devnagari script presented a Devnagari hand-printed numeral recognition system based on binary decision tree classifier. The study investigates the direction of the Devnagari Optical Character Recognition research (DOCR), analyzing the limitations of methodologies for the systems which can be classified based upon two major criteria: the data acquisition process (on-line or off-line) and the text type (machine-printed or hand-written). No matter which class the problem belongs, in general there are five major stages in the DOCR problem: 1. Pre-processing, 2. Segmentation. 3. Feature Extraction, 4.Recognition, 5. Classification The off-line and on-line character recognition techniques have different approaches; they share a lot of common problems and solutions. Since it is relatively more complex and requires more research compared to on-line and machine-printed recognition, off-line handwritten character recognition is selected as a focus of attention Handwriting Recognition Technology has been improving much under the purview of pattern recognition and image processing since a few decades. Hence various soft computing methods involved in other types of pattern and image recognition can as well be used for DOCR. Optical Character Recognition is a process by which we convert printed document or scanned page to ASCII character that a computer can recognize. The document image itself can be either machine printed or handwritten, or the combination of two. Computer system equipped with such an OCR system can improve the speed of input operation and decrease some possible human errors. Recognition of printed characters is itself a challenging problem since there is a variation of the same character due to change of fonts or introduction of different types of noises. Most of the Indian scripts are composed in two dimensions that make them different from Roman script. Therefore, the algorithms developed for Roman script are not directly applicable to Indian scripts. Many works on Indian scripts OCR have been reported. However, none of these works have considered real- life printed text in Devanagari. Consisting of character fusions and noisy environment. The present a complete OCR for printed text that is written in Devanagari script. The OCR has been tested on samples from various magazines and newspapers.

Read the paper · More papers on PaperTik