A Geometric Method for Extracting Images of PDF Files

N Nagajothi, S. S. Dhenakaran, S. Uma Maheswari · 2022 International Conference on Applied Artificial Intelligence and Computing (ICAAIC) · 2022

PDF is a format of the file that is used for showing research articles mostly. PDFs are actually easy to create, but extracting the data from PDF files is a challenging task. Images contain important as well as useful information which cannot be represented by text. The difficult concepts are better explained by an image than the text. In this paper, a feasible way to extract images from the PDF file is explored. This work is attempted to extract images from scholarly research publications of PDF files. The structure of a PDF file is entirely different from other file formats. A geometric method is proposed to get the position of an image in the PDF document and extracted all the images in that document. These images are saved in JPG or PNG or GIF or BMP format as a separate file. The resolution and size of extracted image are the same as the original image and are suitable for printing.

Read the paper · More papers on PaperTik