Dirichlet Process Mixture Model For Document Clustering with Feature Extraction Which Helps In Page Ranking
Nitesh Timande · 2014
To figure out the appropriate number of clusters to which documents should be partitioned is crucial task in document clustering. In this paper, we propose a novel approach, namely DPMFS, which does document clustering and it helps in improving page ranking. Here we are firstly grouping documents into a set of clusters whereas the number of document clusters is determined automatically by the Dirichlet process mixture model secondly identifying the discriminative words and separate them from irrelevant noise words. Our experiment shows that our proposed approach performs well on the synthetic dataset as well as real datasets. The comparison between our approach and some other stage of the art document clustering approach shows that our approach is more robust as well as effective for document clustering.