Research and Implementation of PDF Specific Element Fast Extraction

Shuming Jiang, 李岩 Li Yan · 2023

In order to better understand the enterprise's bidding preference data, a Python platform-based algorithm that automatically recognizes and extracts specific data, extracts the project information in the PDF file of the government procurement contract, and constructs an enterprise portrait of the enterprise participating in government procurement. The algorithm utilizes the built-in library of Python, and our approach improves the algorithm in the open-source library by accurately locating the form pages to improve the parsing efficiency, thus improving the efficiency of information extraction. Designed for different types of government procurement PDF files, identify the text information and form information in the PDF file, the text data for specific content extraction, form information structured and stored in the database. Experiments show that: the algorithm can quickly, accurately, batch extraction of PDF files of specific text content and form information, and converted into structured data deposited into the database, to achieve the desired purpose; its conversion speed and extraction accuracy is significantly faster than the web page conversion tools, etc., and can be batch extraction operations.

Read the paper · More papers on PaperTik