Text clustering using a multiset model

Satoshi Takumi, Sadaaki Miyamoto · 2011

The aim of this paper is to study methods of agglomerative hierarchical clustering which are based on the model of bag of words with text mining applications. In particular, a multiset theoretical model is used and an asymmetric similarity measure is studied in addition to two symmetric similarities. The dendrogram which is the output of hierarchical clustering often has reversals. If we have a reversal, to obtain clusters from the dendrogram becomes difficult. Then, we show the condition that dendrogram have no reversals. It is proved that the proposed methods have no reversals in the dendrograms. Examples based on Twitter and Wikipedia data show how the methods work.

Read the paper · More papers on PaperTik