Header and footer extraction by page association

Xiaofan Lin · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2003

This paper introduces a robust algorithm to extract headers and footers from a variety of electronic documents, such as image files, Adobe PDF files, and files generated from OCR. Compared with the conventional methods based on the page-level layout and format, the proposed strategy considers a page in the context of neighboring pages. Through the page-association, the headers and footers in different patterns can be automatically detected without human interference or individual templates. In addition, fuzzy string match makes the method robust against OCR errors.

Read the paper · More papers on PaperTik