Document Clustering using Compound Words.
Yong Wang, Julia E. Hodges · International Conference on Artificial Intelligence · 2005
Document clustering is a kind of text data mining and organization technique that automatically groups related documents into clusters. Traditionally single words occurring in the documents are identified to determine the similarities among documents. In this work, we investigate using compound words as features for document clustering. Our experimental results demonstrate that using compound words alone cannot improve the performance of clustering system. Promising results are achieved when the compound words are combined with the original single words to be the features. An evaluation of several basic clustering algorithms is also performed in our work for algorithm selection. Although the bisecting K-means method has been proposed as a good document clustering algorithm by other investigators, our experimental results demonstrated that for small datasets, a traditional hierarchical clustering algorithm still achieves the best performance.