Conceptual Relevance Based Document Clustering Using Concept Utility Scale
A. Kousar Nik, K. Subrahmany · Asian Journal of Scientific Research · 2017
Background and Objective: Volume of documents that is generated, processed, stored and retrieved recently is very high and there is integral need of more robust solutions for document retrieval.In this context, clustering is one of the data-mining models to achieve document tracing and retrieval.Many of the document clustering models evinced in contemporary literature, which depends on individual terms of each document as bag of words and clustering documents based on the term frequency, which is a critical constraint of these models.In this paper, the emphasis of this manuscript is to develop a novel document clustering technique that performs clustering by using concept relations of the given documents to achieve more effective document clustering.Methodology: The proposed model tends to cluster the documents based on their concept relations.In order to this, the proposed model depicted a scale called concept utility scale (CUS), which will in use further to identify the concept scope in order to define the clusters from the given document corpora.The feature optimization and document clustering that performed by using t-score, which is a statistical scale for estimating whether the chosen two vectors are similar or diversified.Hereafter, the depicted model denotes as t-CUS.Proposed solution evaluated by renowned statistical metrics like sensitivity, accuracy, fall-out and specificity.Experiments carried out on datasets comprising specific kind of literature.Results: The experimental study evincing that the proposed model is effective in improving the accuracy of clustering and information retrieval.Conclusion: In addition, the computation complexity is low and linear with the proposed solution t-CUS.