Clustering Based on Context Similarity

László Kovács, Tibor Répási, Erika Baksáné Varga, Péter Barabas · 2008

The discovery of word categories is an important step in statistical grammar induction systems. Word categories can be considered as clusters containing words with similar grammatical or semantic behavior. Having a metric space of words, the clustering algorithm will place similar words into the same cluster, whereas dissimilar ones are clustered into different groups. In this paper we propose an approximate word clustering method based on context similarity. The context of a word is defined here as the set of sentences containing the word. The similarity of two words is measured with the similarity of the corresponding context sets. For the calculation of the context-based distance of two words, a hierarchical agglomerative clustering algorithm has been developed, and is presented here.

Read the paper · More papers on PaperTik