NTNU-CORE: Combining strong features for semantic similarity
Erwin Marsi, Hans Jonas Fossum Moen, Lars Bungum, Gleb Sizov, Björn Gambäck, André Lynum · 2013
The paper outlines the work carried out at NTNU as part of the *SEM’13 shared task on Semantic Textual Similarity, using an approach which combines shallow textual, distributional and knowledge-based features by a support vector regression model. Feature sets include (1) aggregated similarity based on named entity recognition with WordNet and Levenshtein distance through the calculation of maximum weighted bipartite graphs; (2) higher order word co-occurrence similarity using a novel method called “Multisense Random Indexing”; (3) deeper semantic relations based on the RelEx semantic dependency relationship extraction system; (4) graph edit-distance on dependency trees; (5) reused features of the TakeLab and DKPro systems from the STS’12 shared task. The NTNU systems obtained 9th place overall (5th best team) and 1st place on the SMT data set. 1