A new approach to short web document creation based on textual and visual information

Martina Zachariášová, Patrik Kamencay, Róbert Hudec, Miroslav Benčo, Slavomír Matúška · 2013

This paper deals with research in area of automatic semantic inclusion of textual and non-textual information of Web documents. The main idea is to create a robust method for extraction of images and textual segments to obtain short web document. Thus, developed method consist of two data types extractions, where both, image and text data extraction are using Document Object Model (DOM) tree. Extracted objects are saved in separated databases followed by the images analysis that defines and describes image object from semantic point of view. Moreover, the semantic descriptions of all modal objects are utilized to short web document creation. We implement our novel method using the Scale Invariant Feature Transform (SIFT) descriptor within a Support Vector Machine (SVM) classifier. Further, in order to obtain a semantic description of objects in static image, the Support Vector Machine (SVM) classification were applied. Finally, semantic inclusion textual and visual information was realized. The developed method has been tested on real and off-line web documents.

Read the paper · More papers on PaperTik