A Comparison Framework of Similarity Metrics Used for Web Access Log Analysis.
Yusuf Yaslan, Zehra Çataltepe · 2007
In this paper, different types of web session similarity metrics are compared and combined for better web session clustering. Syntactic and co-occurrence information are used for similarity calculation. Syntactic information on a web page includes the place of the page in the directory hierarchy. Co-occurrence information is the amount of the occurrences of two web pages in the same sessions. Vector space representation of sessions and cosine, pearson and jaccard similarities between them are also used as a similarity metric. Clustering quality is used as the goodness measure of a similarity. The clustering quality is given by the internal and external cluster similarity. First, clustering quality is evaluated when different similarity metrics (jaccard, pearson, cosine, syntactic, co-occurrence) are used. Similarity calculation is performed for different number of clusters from 1 to the number of sessions. For reasonable cluster numbers (15-100) syntactic similarity results in the best internal cluster similarity, followed by the co-occurrence similarity. Finally the best two methods and others are used for similarity combination. It is found that linear combination of similarity metrics decreases the external cluster similarity. Similarity metrics are also evaluated using Hubert’s statistics. It is found that syntactic similarity alone gives the best results according to Hubert’s statistics followed by the linear combination of syntactic and cooccurrence similarity.