Improving Scene Text Recognition With A Combinative Image Augmentation Approach
Ngan-Linh Nguyen, Gia-Huy Lam, Hoang-Thong Vo, Trong-Hop Do, Anh-Tien Tran, Sungrae Cho · 2022 13th International Conference on Information and Communication Technology Convergence (ICTC) · 2022
Scene text recognition plays an important role in various intelligent systems today such as robotic process automation and self-driving cars. These systems require knowledge of the surrounding scenery, where the words in the scene hold a lot of valuable information. For instance, scene text recognitions can serve the development of smart tourism, smart museums, and self-propelled robots. To increase the practical applicability of the solution, the recognition model needs to meet efficiently both accuracies and processing time. However, before constructing the model, most existing scene text recognition frameworks underrate the importance of a reliable and well-served augmentation stage, all the images in the dataset are often applied with common augmentation functions. In this paper, we present a combinative augmentation framework that takes random augmentation functions from different types into combinations for each individual image. In addition, the framework also randomizes the number of functions taken and every specific parameter for a particular function on the image. Thus, all these tweaks help to greatly increase the pattern diversity in images and pose an improvement in the model evaluation. Evaluated on both seen and unseen test datasets, the framework increased the accuracy by an average of 5,02% on the NRTR model and 2.36% on the VietOCR model.