Unsupervised Estimation of Word Usage Similarity

Marco Lui, Timothy J. Baldwin, Diana McCarthy · 2012

We present a method to estimate word use similarity independent of an external sense inventory. This method utilizes a topicmodelling approach to compute the similarity in usage of a single word across a pair of sentences, and we evaluate our method in terms of its ability to reproduce a humanannotated ranking over sentence pairs. We find that our method outperforms a bag-ofwords baseline, and that for certain words there is very strong correlation between our method and human annotators. We also find that lemma-specific models do not outperform general topic models, despite the fact that results with the general model vary substantially by lemma. We provide a detailed analysis of the result, and identify open issues for future research. 1

Read the paper · More papers on PaperTik