An F-Measure for Evaluation of Unsupervised Clustering with Non-Determined Number of Clusters
Ricard Marxer, P. Purwins · 2008
In unsupervised learning, such as clustering, the problem occurs how to evaluate the results. In particular, neither the number of clusters nor the mapping between eventually known reference classes, e.g. generated from annotations, and the clusters are known. In this report, a method is suggested that adapts the F-measure for supervised classification to the unsupervised case. The task is to group items into k cluster, without knowing k beforehand. If we have labels to these items (not used in the actual clustering), we can evaluate the unsupervised clustering processes. We adapt a measure introduced by [2] that differs from traditional Receiver Operating Characteristics (ROC). ROC measures have often been used in onset detection evaluation and non-incremental clustering. However, in an unsupervised clustering setting, the mapping between the reference classes and the estimated clusters is unknown. Analogously to the confusion matrix in the evaluation of a classification tasks with known mapping, we introduce a mapping matrix that is first constructed by using the onset matching technique presented in [1] adapted to multiple classes. The false positives are treated as matches to an extra empty class in the mapping matrix. Similarly, the false negatives are assigned to an empty cluster. Empty classes and clusters are treated in the same as the other classes and clusters with the exception that their precision and recall do not contribute to the overall precision and recall. Then, the measure considers several hypotheses of class-to-cluster assignments and integrates them by taking averages of their precision and recall measures weighted by each class frequency of occurrence. Let us consider C ground truth classes and K clusters. Be nc,k the number of cooccurrences of class c and cluster k, nc the total number of occurrences of class c. We