Design of Software for Batch Extracting and Merging Document Pages Based on Character Matching
Qiao Sun, Yi Liu, Tong Wu · 2024
In some cases, it is necessary to extract specific pages from a large number of PDF files for subsequent processing. For example, extracting cover and signature pages from a large number of documents makes it easier for people to sign on them. Ordinary PDF file processing software, such as Adobe Acrobat, requires users to operate the software to open each PDF file one by one, browse the entire text, find specific pages, extract and save them, which is extremely time-comsuming. This article uses Python and commonly used PDF processing libraries to design a software for batch PDF specific pages extraction. This software extracts specific pages of multiple files by detecting keywords on the file page, and has an interface that facilitates multiple parameter settings and observation of processing progress.