Layout Based Information Extraction from HTML Documents

Radek Burget · Proceedings of the International Conference on Document Analysis and Recognition · 2007

We propose a method of information extraction from HTML documents based on modelling the visual information in the document. A page segmentation algorithm is used for detecting the document layout and subsequently, the extraction process is based on the analysis of mutual positions of the detected blocks and their visual features. This approach is more robust that the traditional DOM-based methods and it opens new possibilities for the extraction task specification.

Read the paper · More papers on PaperTik