Semi-automated document image clustering and retrieval

Markus Diem, Florian Kleber, Stefan Fiel, Robert Sablatnig · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2013

In this paper a semi-automated document image clustering and retrieval is presented to create links between different documents based on their content. Ideally the initial bundling of shuffled document images can be reproduced to explore large document databases. Structural and textural features, which describe the visual similarity, are extracted and used by experts (e.g. registrars) to interactively cluster the documents with a manually defined feature subset (e.g. checked paper, handwritten). The methods presented allow for the analysis of heterogeneous documents that contain printed and handwritten text and allow for a hierarchically clustering with different feature subsets in different layers.

Read the paper · More papers on PaperTik