Comparison of Visual and Logical Character Segmentation in Tesseract OCR Language Data for Indic Writing Scripts

Jennifer Biggs · 2015

Language data for the Tesseract OCR system currently supports recognition of a number of languages written in Indic writing scripts. An initial study is de-scribed to create comparable data for Tesseract training and evaluation based on two approaches to character segmen-tation of Indic scripts; logical vs. visual. Results indicate further investigation of visual based character segmentation lan-guage data for Tesseract may be warrant-ed.

Read the paper · More papers on PaperTik