A Study of Text Rearrang in Conversion of PDF Documents into HTML

Lin Qin · Computer and Information Technology · 2014

Most of the existing PDF converters fulfill text detection by locating the coordinate of each text element. Specifically,text detection is realized by rearranging these elements from left to right as well as from top to bottom in the order.Unfortunately, such methods fail to work in complex multiple-column PDF documents. To settle this problem, this work proposed a novel page segmentation algorithm. The proposed algorithm first divides a page into several blocks, and then reorders these blocks. With the proposed algorithm, the correctness of returning to original complex multiple-column text increases effectively.

Read the paper · More papers on PaperTik