Multi-Layer Semantics Based Document Clustering
Muhammad Rafi, Muhammad Sharif, Waleed Arshad, Sheharyar Mohsin, Habibullah Rafay · 2016
Document clustering is an unsupervised machine learning method that separates a large subject heterogeneous collection (Document Base or Corpus) into smaller, more manageable subject homogeneous collections (clusters). Traditional method of document clustering uses features like words, sequence, phrases, etc. These features are independent to each other and do not cater semantics. In order to perform semantic viable clustering, we believe that the problem of document clustering has two main components: (1) to represent the document in such a form that it inherently captures semantics of the text. This may also help to reduce dimensionality of the document and (2) to define a similarity measure based on the semantic representation such that it assigns higher numerical values to document pairs which have higher semantic relationship. In this paper, we propose a representation of document, based on three distinct layers: these are lexical, syntactic and semantic layers. We believe that these three layers are essential to ensure semantics into document meta descriptor. Using these three layer's features we propose a similarity function for performing document clustering. We performed an extensive series of experiments on standard text mining data sets with external clustering evaluations like: F-Measure and Purity.