Algorithms for extracting lines, paragraphs with their properties in PDF documents

Viacheslav Martsinkevich, Andrei Berezhkov, Vladislav Tereshchenko, Natalia Nikolaevna Gorlushkina, Violetta Tretjakova · E3S Web of Conferences · 2023

The article discusses the algorithms for detecting and extracting lines, paragraphs with their properties and attributes in PDF documents, analyses the structure of PDF-file and its objects. Due to special operators in objects the PDF documents content is saved as symbols or symbol groups. The position of such groups on the page also remains identical. The main challenge that we face, while extracting paragraphs from the PDF document is the complex format that is able to retain various types of information and can be created in several ways.

Read the paper · More papers on PaperTik