Identifying contents page of documents

Qin Luo, Takashi Watanabe, T Nakayama · 1996

Contents pages of a document include useful information such as the list of contents, the hierarchical organization of the document, and the locations of each component (page number). Therefore, identification of contents pages is an effective way to establish the reference information of documents automatically. In this paper, we propose a recognition method for contents pages of documents: not only extract the meaningful data from contents page images and classify them into distinct items, but distinguish contents pages from other pages. As many researches reported until today indicate, successful methods for document analysis depend on the effective application of object-specific information. In this sense, the recognition of document types or classes is an important complement to document analysis.

Read the paper · More papers on PaperTik