Constraint Word Clustering Algorithm for Asymmetric Relationship

Rajkumar Jain, Narendra S. Chaudhari · 2015

A new constraint word clustering algorithm is proposed for the given corpus. The proposed method is based on the constraint clustering of words. In this paradigm words are considered similar if they appear in similar contexts and contexts are similar if there word affinity clouds are equivalent. Different types of association between words are identified and based on this association constraints are identified and generated. Proposed constraint algorithm is applicable for words having asymmetric relationship between them; therefore this approach may prove useful as a complement to conventional class-based statistical language modeling techniques. In the context of information retrieval, a new constraint word clustering is projected based on the paradigm of constraints for asymmetric relationship between words. Constraint word clustering approach is appropriate at discovering semantic relationship between words rather than discovering syntactic relationship between words. Affinity (3, 4, 5) describes the quantitative relationship between words. An affinity describes a quantitative relationship between the two words and this in turn helps to identify the clusters of words. A cluster comprises words that are sufficiently affine with each other. A first word is sufficiently affine with a second word if the affinity between the first word and second word satisfies one or more affinity criteria. Present research focuses on the clustering of words based on the finding of semantic relationship between words. Semantic relationships between words are modelled by identifying the constraints. Present research proposes a constraint clustering architecture and algorithm based on the different types of constraint associated between words. Our contribution is summarized as follows: we investigated the constraints based on properties word cloud, we investigated the constraints based on association between words and we presented a constraint word clustering algorithm for asymmetric relationship between words.

Read the paper · More papers on PaperTik