Enhancing the Searchability of Page-Image PDF Documents Using an Aligned Hidden Layer from a Truth Text

Ian A. Knight, David F. Brailsford · 2016

The search accuracy achieved in a PDF image-plus-hidden-text (PDF-IT) document depends upon the accuracy of the optical character recognition (OCR) process that produced the searchable hidden text layer. In many cases recognising words in a blurred area of a PDF page image may exceed the capabilities of an OCR engine.

Read the paper · More papers on PaperTik