Localized document image change detection

Rajiv Ratan Jain, David Doermann · 2015

Given two versions of a document image, the goal of document image change detection is to automatically determine exactly what content was added, deleted or modified. Typically, one would accomplish this by first performing Optical Character Recognition (OCR) on the two documents and then performing a “diff” to identify the changes. However, this approach can fail due to OCR errors, poor segmentation, or the inability to handle graphical content. We compare the OCR baseline with two techniques based on SIFT features that detect changes in the image at the word level. The first approach performs the “diff” on SIFT features extracted from the center line of the text image. The second approach performs a segmentation free alignment of text blocks using dense SIFT to address the more general cases where segmentation fails or graphical objects are modified. Results on two experimental datasets show the improvement of the segmentation free approach over the baseline approach.

Read the paper · More papers on PaperTik