Document Clustering using Weighting and Labels based on Inherent Structure of Document

Yong‐Il Kim, Yoo-Kang Ji, Sun Park · 2013

In classic document clustering, documents appear terms frequency without considering the semantic information of each document (i.e., vector model). The property of vector model may be incorrectly classified documents into different clusters when documents of same cluster lack the shared terms. Recently, to overcome this problem uses knowledge based approaches. However, these approaches have an influence of inherent structure of documents on clustering and a cost problem of constructing ontology. This paper proposes a new document clustering method using terms of class label and term weights based on inherent structure of documents by NMF. The experimental results demonstrate that the proposed method achieves better performance than other document clustering methods.

Read the paper · More papers on PaperTik