Automatic Websites Classification and Retrieval using Websites Communication Signatures

Jean-Charles Lamirel, David Reymond · Collnet Journal of Scientometrics and Information Management · 2014

Automated classification and summarization of websites, as well as knowledge retrieval from web contents, are central challenges for performing accurate and focused webometrics studies. As global approaches based on open web and full webpages content fail to cope with such challenges, in this paper we first focus our approach on organizational and institutional websites, and secondly, we consider the communicational value of the data provided, especially on navigation menus, as a central information source. Another key point of our approach is that we more especially focus on the exploitation of a recent unsupervised classification technique, that not only automatically groups together websites sharing a number of features but also explicitly associates each website class with a set of specific features, or labels, characteristic of that class. As compared to a supervised classification, our approach presents the main advantage to cope with the scaling problem whilst, as we show in our experiment, providing superior performance through efficient characterization of latent unlabeled websites classes.

Read the paper · More papers on PaperTik