Hybrid Edge Detection-Based Approach for Enhanced Text Extraction from Image
Akash Burnwal, Mamata P. Wagh, Nishi Choudhary, Saurav Kumar · 2025
Text extraction from images plays a crucial role in applications such as document digitization, automated data extraction, and computer vision-based text processing. This paper introduces a multi-stage pipeline to enhance text recognition accuracy through preprocessing, Hybrid Edge Detection (HED), and Optical Character Recognition (OCR). Preprocessing removes noise using Gaussian blurring, followed by grayscale conversion to improve contrast. The HED phase combines Canny and Sobel techniques to detect text contours, generating a hybrid edge map by applying an OR operation on the outputs of both methods, preserving all detected edges. This fused edge map is then processed through OCR to extract character patterns. In scenarios where external noise was deliberately introduced to test the robustness of various approaches, traditional methods failed to detect text effectively, whereas the proposed model demonstrated resilience in adverse conditions. The proposed HED scheme enhances the accuracy of text extraction, achieving the highest Word Match Accuracy of 83.3%.