Opportunities and Challenges in Semantic Similarity
Vasile Rus · The Florida AI Research Society · 2014
Semantic similarity has been increasingly adopted in the recent past as a viable, scalable alternative to the full-understanding approach to natural language understanding. We present here an overview of opportunities and challenges in semantic similarity research with an emphasis on methods, data, and tools. A series of methods we developed over the past decade will be summarized. These methods and others have been integrated in a semantic similarity toolkit called SEMILAR ( www.semanticsimilarity.org ), which has been widely adopted by thousands of users sinces its launch in summer of 2013 at the Annual Meeting of the Association of Computational Linguistics. Furthermore, we illustrate some drawbacks of current data sets that hamper a fair comparison of existing methods. Several suggestions will be made to improve the building of future data sets for assessing the performances of approaches to semantic similarity.