Extracting Images from Chinese PDF Documents

Yong Hua Yin, Ying Jin, Quanyin Zhu, Yun Yang Yan · Applied Mechanics and Materials · 2014

In order to efficient tap the potential value in Chinese PDF documents and use Chinese PDF documents, an unique idea that extracting images from Chinese PDF documents is proposed in this paper. The idea combines PDFs document structure and page tree to extract images. Based on this idea, the experiments in this paper are done with one hundred Chinese PDF documents. And the extraction rate of the experiments obtains 83.56 percent. According to the analysis of experimental results, it is proved that the idea proposed in this paper is applicable to most of Chinese PDF documents and it is able to meet most of the needs of practical application.

Read the paper · More papers on PaperTik