Resilient Kannada Scene Text Detection: CRAFT-YOLOv8 Fusion

Anoshor B Paul, G Sai Vikrant, V. Umadevi · 2024

The extraction of text from natural scene images is an increasingly vital task in computer vision, with applications spanning numerous domains. This study introduces a simple and effective method tailored for detecting Kannada text in complex scene images, extendable to use in any Indian language. This paper presents a novel approach for Kannada scene text detection that combines the CRAFT (Character Region Awareness for Text Detection) algorithm for text annotation generation with the YOLOv8 (You Only Look Once) model for text detection. The paper details a simple implementation of the aforementioned algorithm, addressing challenges such as varying text sizes, orientations, and complex backgrounds. The dataset annotation process was streamlined and optimized, saving significant time compared to traditional methods. To compensate for the scarcity of Kannada text datasets, a custom Kannada scene image dataset was curated. This dataset offers diverse real-world scenarios for testing the model. The model achieved an accuracy of 85.30% mAP at a 0.6 Intersection Over Union (IoU) threshold on the training-validation dataset and a commendable 80.10% mAP across a diverse set of Kannada test images. The primary objective of this paper is to furnish researchers with an accessible methodology for detecting local Indian languages in scene images and to further extend this research to the development of more sophisticated text extraction tools, benefiting a diverse range of digital applications.

Read the paper · More papers on PaperTik