Scan.it - Text Recognition, Translation and Conversion

Minal Acharya, Priti Chouhan, Asmita Deshmukh · 2019

Electronic documents are in demand everywhere in workplace, government offices, schools, colleges. There are many texts - such as letters and documents- which are available in electronic form but need to be converted in some readable regional language. However, conventional scanners are limited by their large size and relatively cumbersome usage. In this paper we explore the need for document scanning, and how this portable application can help users overcome the language barrier among states based on regional language by converting scanned documents. Scan.it web application is a Natural Language Processing based system which is used for text recognition and translation in an intelligent way. Scan.it solves the problem of ambiguity between similar words in a text and use a more semantic approach so as to classify data according to the context it is used in. The code can be implemented using Node.js. Further Tesseract OCR gives the results of the recognized text and ImTranslator is incorporated to translate this recognized text. In contrast to other OCR techniques, Scan.it solves the problem of recognizing the Marathi language text and further also helps in maintaining the soft copy of the document.

Read the paper · More papers on PaperTik