Semantically enriching text representation model for document clustering
Hanjoon Kim, Kee-Joo Hong, Jae Young Chang · 2015
This paper presents a novel text space model that represents textual documents for document clustering, which contains the 'concept' space independently of the 'document' and 'term' spaces. The text model described here represents documents as matrices (i.e., 2nd-order tensors), and a document corpus is represented as a 3rd-order tensor. For this, it is necessary to produce the concept vector for each term that occurs in a given document, which is related to word sense disambiguation. As an external knowledge source for concept weighting, we employ the Wikipedia encyclopedia.