Multimedia information extraction from HTML product catalogues

Martin Labský, Pavel Praks, Vojtěch Svátek, Ondřej Šváb · DATESO · 2005

We describe a demo application of information extraction from company websites, focusing on bicycle product offers. A statistical approach (Hidden Markov Models) is used in combination with different ways of image classification, including latent semantic analysis of image collections. Ontological knowledge is used to group the extracted items into structured objects. The results are stored in an RDF repository and made available for structured search.

Read the paper · More papers on PaperTik