An Algorithmic Approach for Text Recognition from Printed/Typed Text Images

Neha Agrawal, Arashdeep Kaur · 2018

Extraction of texts from scanned copies of documents and text images is an important task in the recent scenario. Optical Character Recognition (OCR) is used to analyze text in images. The proposed algorithm deals with taking scanned copy of a document as an input and extract texts from the image into a text format using Otsu's algorithm for segmentation and Hough transform method for skew detection. The system was confined to recognize English alphabets (A-Z, a-z) and numerals (0–9). OCR technique has been implemented to recognize characters. Validation tests were done on screenshots of typed texts and images of scanned document from Internet sources. Experimental results indicate that the proposed algorithm is able to recognize alphabets written in Verdana font style with size 14 and also showed good results with rotated images. The average accuracy to determine rotation angle correctly was calculated to be 90% and overall system accuracy was calculated to be 93%.

Read the paper · More papers on PaperTik