Semantic Oriented Document Clustering Using Distribution Semantics
Umar Ali Khan, Muhammad Rafi · 2018
The exponential growth of electronic form of textual documents in both public and proprietary storage force researchers to find way to efficiently extract meaningful, actionable information from these documents. Document clustering has find its niche in this area. This paper proposes a document representational model based on distributional semantics, the law of distributional semantics says that the linguist terms that appear with similar distribution in a language corpus generally have similar meaning. This representation of document model uses only those terms (linguistic feature) that have same distribution over a given collection of documents. So to find this, it is needed to find out the distributional terms by using distributional criteria and then representing the documents by only these distributional terms. A novel similarity measure is proposed over these documents that also utilized the very nature of distributional semantics in similarity calculation. Finally, hierarchal agglomerative clustering (HAC) is used to produce the final clusters. Standard text mining datasets are used to measure the effectiveness of this approach. The evaluation is based on purity of clusters and proposed approach achieved far better clustering results in comparison to conventional approach.