Incremental hierarchical clustering of text documents

Nachiketa Sahoo, Jamie Callan, Ramayya Krishnan, George T. Duncan, Rema Padman · 2006

A version of cobweb/classit is proposed to incrementally cluster text documents into cluster hierarchies. The modification to classit consists of changes to the underlying distributional assumption of the original algorithm that are suggested by text document data. Both the algorithms are evaluated using standard text document datasets. We show that the modified algorithm performs better than the original Classit when presented with Reuters newswire articles in temporal order, i.e., the order in which they are going to be presented in real life situation. It also performs better than the original Classit on the larger of eleven standard text clustering datasets we used.

Read the paper · More papers on PaperTik