Optical Character Recognition: Comparative Analysis of Tesseract and Textract on Diverse Datasets
Manan Modi, Arpita Shah, Deep Kothadiya, Mrugendrasinh L. Rahevar · 2024
Optical Character Recognition (OCR) is essential for converting images and scanned documents into machine-readable text. Advances in deep learning and machine learning have enhanced OCR capabilities, allowing for tasks such as handwritten text recognition and structured data extraction. This paper compares two OCR engines: Tesseract, an open-source tool, and Amazon Textract, a cloud-based service. Their performance is tested on three datasets; printed text, handwritten documents, and financial records. Amazon Textract demonstrates superior accuracy and scalability due to its cloud infrastructure, while Tesseract remains a cost-effective option but struggles with complex layouts. This study offers insights for selecting OCR tools based on document complexity and processing needs.