Unsupervised extraction of product information from semi-structured sources

Maximilian Walther · 2012

Product information search has become one of the most important application areas of the Web. Especially considering pricey technical products, consumers tend to carry out intensive research activities previous to an actual acquisition. However, the vast amount of available data about such products and its various representations may easily overstrain potential customers. In this paper, we develop a comprehensive technique for extracting product specifications about arbitrary technical products from web pages in a widely unsupervised manner. The technique is based on a clustering approach that uses structural and visual features of web page elements. The resulting detailed information sets allow a potential consumer to effectively compare products while saving the manual extraction work.

Read the paper · More papers on PaperTik