Text Recognition from Images using a Deep Learning Model
Sanda Sri Harsha, B P N Madhu Kumar, R.S.S. Raju Battula, P. John Augustine, S. Sudha, T Divya. · 2022 Sixth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC) · 2022
Identifying text in an image or video is a process called text recognition. Both documents and other real-time visuals, including arbitrary photos, use this method. For this, numerous hardware elements and algorithms were created. This study tries to determine which of the two algorithms has the best feature extraction method. A dataset of arbitrary real-time photos with text is collected for this use using an ICDAR dataset. Every image in the collection is preprocessed using one of two techniques. They are image scaling and picture enhancement. The preprocessed photos are then asked for text localization. The cropped image is then examined with character-level grouping. Two feature extraction algorithms are used to extract the key features. The extraction techniques are the HOG or the Histogram of Oriented Gradients method and the LBP or the Local Binary Pattern recognition technique. A deep learning model is developed using the convolutional neural network method. The model then recognizes the term. For each methodology, all the stages are repeated twice to identify the best extraction techniques. The optimum feature extraction method is then determined by putting the models to the test. The results of the models are based on three factors. They are accuracy, precision, and the F1 score. In the end, it is found that the LBP algorithm is better than the HOG algorithm in all three parameters.