DEDUCTION TECHNIQUES FOR UNSUPERVISED CLUSTERING WITH SIMILARITY HIDDEN CLUSTERING MODEL
N. Senthilkumaran, M. Phil, S. Sesha Sai Priya · 2015
Document clustering is automatically group related document into cluster. In this clustering frame work focus on correlations between the documents in the local patches are maximized while the correlations between the documents outside these patches are minimized simultaneously. The proposed systems are adopts both supervised and unsupervised constraints to demonstrate the effectiveness of the proposed algorithm in this framework. The novel proposed classical K-Mean clustering algorithm applied for data preprocessing in stop word removal, stemming and synonym word replacement to apply semantic similarity between words in the documents. In addition, content can be retrieved from text files, HTML pages as well as XML pages. Tags are eliminated from HTML files. Attribute name and values are taken as normal paragraph words in XML files and then preprocessing (stop word removal, stemming and synonym word replacement) is applied. In addition, TEXT, HTML and XML documents are cluster using cosine similarity model implement these research works.