Unsupervised extraction of product information from semi-structured sources
Maximilian Walther · 2012
Product information search has become one of the most important application areas of the Web. Especially considering pricey technical products, consumers tend to carry out intensive research activities previous to an actual acquisition. However, the vast amount of available data about such products and its various representations may easily overstrain potential customers. In this paper, we develop a comprehensive technique for extracting product specifications about arbitrary technical products from web pages in a widely unsupervised manner. The technique is based on a clustering approach that uses structural and visual features of web page elements. The resulting detailed information sets allow a potential consumer to effectively compare products while saving the manual extraction work.