Benchmarking Performance Analysis of Optical Character Recognition Techniques

Muhammad Ali Naqi Hadi, Maria Gul, Maqbool Khan, Ghadah Naif Alwakid, Noor Zaman Jhanjhi · 2024

Optical Character Recognition automates the extraction of printed and handwritten text from documents; it thus is very vital in digitalizing records. This research benchmarks seven optical character recognition (OCR) engines: Paddle-OCR, Easy-OCR, Keras-OCR, Pytesseract OCR, OpenCV OCR, PyMuPDF OCR, and DocTR OCR, on 200 diverse CBC patient reports. The metrics taken into consideration for evaluation included execution time, accuracy, Character Error Rate, and Word Error Rate. Finally, PaddleOCR is the one that performs best: 67.28% in accuracy, 0.43 in character error rate (CER), and 0.66 in word error rate (WER). Meanwhile, Keras OCR records the highest error rates among all. The findings make an insight into picking an appropriate OCR tool, which is bound to have tremendous implications for healthcare, legal documentation, and historical preservation in guiding future OCR advancements.This research will be baseline for further research in OCR.

Read the paper · More papers on PaperTik