Robust, Light-weight Approaches to compute Lexical Similarity
Quang Việt Ðỗ, Dan Roth, Mark Sammons, Yuancheng Tu, V. G. Vinod Vydiswaran · 2010
Most text processing systems need to compare lexical units – words, entities, semantic concepts – with each other as a basic processing step within large and complex systems. A significant amount of research has taken place in formu-lating and evaluating multiple similar-ity metrics, primarily between words. Often, such techniques are resource-intensive or are applicable only to spe-cific use cases. In this technical re-port, we summarize some of our re-search work in finding robust, light-weight approaches to compute similar-ity between two spans of text. We describe two new measures to com-pute similarity, WNSim for word similar-ity, and NESim for named entity sim-ilarity, which in our experience have been more useful than more standard similarity metrics. We also present a technique, Lexical Level Matching (LLM), to combine such token-level similarity measures to compute phrase-and sentence-level similarity scores. We have found LLM to be useful in a number of NLP applications; it is easy to compute, and surprisingly robust to noise. 1