Extraction of PDF Table Data Based on the Pdfplumber Method
Wen Zhi Yang, Feifei Cao, Xueli Zhao · 2024
A large amount of information is stored in various electronic documents, including scientific literature, reports, and research papers. The table in these documents is one of the main ways to contain rich and critical information. To edit and reuse the table data in PDF files, automatic recognition and extraction of PDF table data is implemented based on the pdfplumber parsing method. First, the pdfplumber converts PDF files into the parsable Python objects, including text, images, tables, and other information in the file. Then it parses the specified PDF pages in the Python object, converts them into the table data, and automatically extracts them into Excel files. The experimental results show that the integrated process can identify and extract the table data in PDF files in batches. Using the recognition method of pdfplumber, the recognition rate of table data such as the survey tables and the education resource tables in higher education quality reports reaches an average of 96%, effectively handling complex situations such as blank spaces and column misalignment in the tables.