An approach for Text Recognition from Document Images

Sujata S. Desai, Darshana Rajput, Kiran Patil · 2020 IEEE Bangalore Humanitarian Technology Conference (B-HTC) · 2020

In the recent scenario, it is important to extract text from scanned documents and text images. Analyzing of the text within the images utilizes the optical character recognition (OCR). The proposed system deals with the utilization of the Otsu segmentation algorithm and Hough transforming method of bias detection within the kind of input and separating the content from the photographs into a word format. In this system, the alphabets plus the numerals of English were known. In order to recognize the characters, OCR method has been applied. The screenshots of the scanned documents to the photographs from the searched document on the web resources performed validation checks. Experimental results demonstrate that the alphabets written in Verdana fonts of 14 size is recognized by the proposed system and even have got great outcomes together with pivoted images. The efficiency for correctly determining the rotatory angle was calculated at 90% and also the total accuracy of the device was calculated at 93 percentage.

Read the paper · More papers on PaperTik