Optical Character Recogniation for Tigrigna Printed Documents Using HOG and SVM

Kibrom Tsegay Fitsum, Yask Patel · 2018

Optical character recognition (also OCR) is the process of converting and recognizes characters from image to digital format (ASCII). Tigrigna language is one of Ethiopian language which is spoken in the northern part of Ethiopia. There are a lot of printed documents which is written in Tigrigna language. This work presents an Optical Character Recognition (OCR) to convert printed documents into computer format using SVM classifiers. The input image convert into Grayscale and binary (black and white) and then remove noise and skew detection and correction. Segmentation is the most significant matter in order to segment the character by word, line, and character. Feature extraction used to extract information from the character. We are using HOG (histogram of oriented gradient) techniques to extract the features of each segmented images and SVM classifier used to classify based on the feature extraction in order to train the machine. The huge amount of documents piled high in information centers, libraries, and government and private offices in the form of correspondence letters, magazines, newspapers, pamphlets, books, etc. Converting these documents into electronic format is a must in order to preserve historical documents, save storage space and Enhance retrieval of relevant information via the Internet. This enables to harness existing information technologies to local information needs and developments.

Read the paper · More papers on PaperTik