Unsupervised Clustering of Images using their Joint Segmentation
Yevgeny Seldin, S Starik, Michael Werman · 2003
We present a method for unsupervised content based classification of images. The idea is to first segment the images using centroid models common to all the images in the set and then, through drawing an analogy between models/images and words/documents, to apply algorithms from the field of unsupervised document classification to cluster the images. The first step may be regarded as unsupervised feature selection while the second may be regarded as unsupervised classification of images based on the selected features. We regard our image set as a mixture of textures. The centroid models of the mixture representing the textures are based on histograms of marginal distributions of wavelet coefficients calculated on image subwindows. The models are used in our algorithm (which is analogous to the work of Hofmann, Puzicha and Buhmann [HPB98]) to jointly segment all the images in the input set. Such joint segmentation enables us to link between multiple appearances of the same texture in different images. We finally use the sequential Information Bottleneck algorithm of Slonim, Friedman and Tishby [SFT02] to cluster the images based on the result of the segmentation. In general, due to the modularity of the approach each of the three components of the presented method (local image modeling, segmentation and classification) can be substituted by alternative algorithms satisfying mild conditions. The method is applied to nature views classification and painting categorization by painter’s drawing style. The method is shown to be superior to image 1 classification algorithms that regard each image as a single model. We see our current work as opening a new perspective on high level unsupervised data analysis. 1.