TOC generation in PDF Document for Smart Automated Compliance Engine
Muzammil Hussain Shahid, Muhammad Arshad Islam · 2020
Portable Document Format (PDF) is a commonly used format for the scientific publication. Currently, an input document is used to test the compliance and relevance of the document or text in Automated Compliance Engines and Natural Language Processing(NLP) based system. The whole document text is used for searching the compliance rules which is computationally expensive and slow process. For speeding up the compliance checking process and making it cost efficient, this paper purposes a method based on Table of Content(TOC) Data Structure. This work proposed the PDFparser which performs Data Indexing, separate headings text, and non-heading text, create hierarchy of headings and generates TOC to reduce the semantic-based string searching time and space. Furthermore, in the NLP based system, mostly semantic-based string matching used. The proposed PDFparser uses the Cosine Similarity method for computing semantic based similarity. Our purposed method performs 47.2% better than the previous approach of searching in the non-indexed whole document and decreases the search time and space. In the worst-case scenario, where no string match found, our purposed method performs 20.5 % better.