Optical Character Translation Using Spectacles (OCTS)
Kanagasabai Thiruthanigesan, Roshan Ragel · 2019
Optical Character Recognition (OCR) is reproducing the text as a digital format that has been produced by the non-computerized system. The translation is an essential part because people have barriers in languages. Especially when people read articles and books in their native language, they can understand more clearly and can get more ideas about the content. The framework and the portable hardware system developed takes images of printed documents and converts images as OCR text by using tesseract and then translates the text by using Google Translator API. The framework implements image capturing techniques, optical character recognition and translation using an embedded system based on Raspberry Pi. Daily used sentences were selected and captured in different development stages, viewpoints, angles, and background. Then the images are classified through the proposed system. Several pre-processing mechanisms were introduced in the process of OCR to get quality output via improving the accuracy rate of the system. From the experiment carried out on one hundred and twenty-three daily used sentences, 97% of the sentences were correctly detected and 91% Tamil and 89% Sinhala sentences were correctly translated.