Document Segmentation and Language Translation Using Tesseract-OCR

Thakare Sahil, Ajay Kamble, Vishal Thengne, Uday Kamble · 2018

Document segmentation and Translation are one of the key areas in pattern recognition and natural language processing. This paper presents details about translation in terms of a web application that accepts image document as an input, where input document is a user define image file containing text in any language available in the Python-tesseract library and does its exact translation in any supported languages using Google Translator (i.e Googletrans). Python script and various libraries are used to approach various challenges in segmentation and translation of a document.

Read the paper · More papers on PaperTik