Tree-Based Structural Representation and Difference Computation for PDF Documents

Nancy Thomas, Daniel Borrajo · Research Square · 2023

Abstract Purpose: The goal of this paper is to introduce a method that uses the structured representation of two PDF documents to compute differences between them. Methods: Included in this work there are three main contributions. The first one is a method for parsing PDF documents and representing them with a tree structure. The second one is our algorithm for computing differences between two PDF documents by leveraging their tree structure. The third one is a method for generating pairs of original and modified synthetic PDF documents. Results: Our system for parsing and representing PDF documents is able to detect the true structure of a PDF document with high accuracy when tested on synthetic PDFs. Our tree-based difference computation approach, when tested on synthetic documents, is able to capture structural differences and improve computational complexity significantly compared to the performance of state of the art alternative approaches. Conclusion: Tree-based PDF document representations allow for efficient, accurate, and comprehensive computation of differences between documents.

Read the paper · More papers on PaperTik