Text Clustering using Semantics
Bhoopesh Choudhary, Pushpak Bhattacharyya · 2002
In traditional document clustering methods, a document is considered a bag of words. The fact that the words may be semantically relateda crucial information for clusteringis not taken into account. In this paper we describe a new method for generating feature vectors, using the semantic relations between the words in a sentence. The semantic relations are captured by the Universal Networking Language (UNL), which is a recently proposed semantic representation for sentences. The clustering method applied to the feature vectors is the Kohonen Self Organizing Maps (SOM). This is a neural network based technique, which takes the vectors as inputs and forms a document map in which similar documents are mapped to the same or nearby neurons. Experiments show that if we use the UNL method for feature vector generation, clustering tends to perform better than when the term frequency based method is used.