Automatic verification of the text layer correctness in PDF documents

Oksana Vladimirovna Belyaeva, Aleksandr Golodkov, Bekzat Bukhatov · 2024

PDF documents can contain incorrect textual layers due to low scanning quality, font embedding errors or other reasons. An incorrect text layer can significantly hamper automatic document processing and limit document parsing. To help tackle the problem, this work analyzed the causes of incorrect text layers (gibberish characters) and described a method to automatically extract text information by detecting the correct text layer of PDF files, using OCR and a PDF extractor. To determine the correctness/incorrectness of the text layer obtained from PDF extractors in the method, we trained a binary classifier on the generated data.We carried out experiments and assessed the quality of various classification models for text correctness, tested the method on real data, thereby confirming the effectiveness of the developed approach. The obtained results can be useful for further development of methods for automatically checking textual layers and improving the quality of automatic processing PDF documents.

Read the paper · More papers on PaperTik