Concept Based Document Clustering Using K Prototype Algorithm
Sneha Pasarate, Rajashree Shedge · 2018 International Conference on Control, Power, Communication and Computing Technologies (ICCPCCT) · 2018
Internet is a wide network of unstructured data such as blogs, tweets, mails, files. The retrieval of correct data is necessary. Data clustering is the process of putting together data into groups which are coherently similar. Clustering minimizes search time as similar documents are in the same cluster. Named entity recognition method allocates noun entities to different sections named as person, location etc. Latent dirichlet allocation treats documents as mixture of topics, and works as generative model. We have used Reuters 21578 dataset, self created dataset, News article dataset, web dataset for processing. The proposed system consists of preprocessing web document data for removing unwanted data. Next is the feature extraction phase through named entity recognition method and topic modeling approach(LDA). Feature extraction shrinks data dimensionality. K-prototype clustering algorithm approach performs better for clustering as it takes into consideration number of mismatches for categorical data. The execution time and space utilized by K-prototype algorithm is better than Fuzzy clustering algorithm.