Suitability of OCR Engines in Information Extraction Systems : a Comparative Evaluation
Zacharias Erlandsson · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2019
Previous research has compared the performance of OCR (optical character recognition) engines strictly for character recognition purposes. However, comparisons of OCR engines and their suitability as an intermediate tool for information extraction systems has not previously been examined thoroughly. This thesis compares the two popular OCR engines Tesseract OCR and Google Cloud Vision for use in an information extraction system for automatic extraction of data from a financial PDF document. It also highlights findings regarding the most important features of an OCR engine for use in an information extraction system, when it comes to structure of output as well as accuracy of recognitions. The results show a statistically signifant increase in accuracy for the Tesseract implementation compared to the Google Cloud Vision one, despite previous research showing that Google Cloud Vision outperforms Tesseract in terms of accuracy. This was accredited to Tesseract producing more predictable output in terms of structure, as well as the nature of the document which allowed for smaller OCR processing mistakes to be corrected during the extraction stage. The extraction system makes use of the aforementioned OCR correctional procedures as well as an ad-hoc type system based on the nature of the document and its fields in order to further increase the accuracy of the holistic system. Results for each of the extraction modes for each OCR engine are presented in terms of average accuracy across the test suite consisting of 115 documents.