Implication intensity: Randomized F-measure for cluster evaluation
Limin Li, Junjie Wu, Shiwei Zhu · 2009
The ever-growing resources of information and services on World Wide Web provide a welcome boost for the researches in the information retrieval space. Text clustering groups a set of documents into subsets or clusters so that the vast retrieved documents can be browsed selectively and efficiently. Many cluster validation measures, such as the F-measure, are then introduced to evaluate the clustering qualities. In this paper, however, we demonstrate that this widely adopted F-measure suffers from the so-call increment effect which may mislead the comparison of clustering results with different cluster numbers. To meet this challenge, we propose a novel ldquoimplication intensityrdquo (IMI) measure based on the F-measure and a random clustering perspective. Experimental results on real-world data sets demonstrate that IMI shows merits on alleviating the increment effect introduced by the F-measure.