Comparison of Visual and Logical Character Segmentation in Tesseract OCR Language Data for Indic Writing Scripts
Jennifer Biggs · 2015
Language data for the Tesseract OCR system currently supports recognition of a number of languages written in Indic writing scripts. An initial study is de-scribed to create comparable data for Tesseract training and evaluation based on two approaches to character segmen-tation of Indic scripts; logical vs. visual. Results indicate further investigation of visual based character segmentation lan-guage data for Tesseract may be warrant-ed.