Non-Manhattan layout extraction algorithm
Aziza Satkhozhina, Ildus Ahmadullin, Jan P. Allebach, Qian Lin, Jerry Liu, Daniel R. Tretter, Eamonn O’Brien-Strain, Andrew Hunter · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2013
Automated publishing requires large databases containing document page layout templates. The number of layout templates that need to be created and stored grows exponentially with the complexity of the document layouts. A better approach for automated publishing is to reuse layout templates of existing documents for the generation of new documents. In this paper, we present an algorithm for template extraction from a docu- ment page image. We use the cost-optimized segmentation algorithm (COS) to segment the image, and Voronoi decomposition to cluster the text regions. Then, we create a block image where each block represents a homo- geneous region of the document page. We construct a geometrical tree that describes the hierarchical structure of the document page. We also implement a font recognition algorithm to analyze the font of each text region. We present a detailed description of the algorithm and our preliminary results.