Character Segmentation Technique for Printed and Handwritten Devanagari Script without Extraction of Shirorekha
Ambadas B. Shinde, Yogesh Hari Dandawate · ADBU - journal of engineering and technology · 2020
In India, there is a lot of literature available in the Devanagari script as well as Devanagari is most frequently used for written, oral correspondence and documentation reasons. How correctly the character segmentation of the Devanagari content is done, will ultimately decide the exactness of the OCR process. In this paper, we have proposed the character segmentation strategy along with existence of Shirorekha framed for printed as well as handwritten text written in Marathi language. Several methods utilized for pre-processing the document images like document binarization, skew identification and correction are discussed in this paper. We took vertical projection of the segmented words, compared the pixel count with the automatically calculated threshold and then characters were separated. With the proposed strategy, we have achieved 100 % exactness in line and word segmentation and results accomplished for character segmentation are a lot of encouraging. No standard dataset of Marathi characters is available having upper and lower modifiers. Non-availability of the standard datasets with modifiers is a significant issue in using a deep learning network for perceiving the Devanagari characters. The segmented characters with the presence of Shirorekha can be straightforwardly utilized for developing the deep learning OCR framework.