ECNU: Using Traditional Similarity Measurements and Word Embedding for Semantic Textual Similarity Estimation

Jiang Bo Zhao, Man Lan, Jun Feng Tian · 2015

This paper reports our submissions to seman-tic textual similarity task, i.e., task 2 in Se-mantic Evaluation 2015. We built our sys-tems using various traditional features, such as string-based, corpus-based and syntactic simi-larity metrics, as well as novel similarity mea-sures based on distributed word representa-tions, which were trained using deep learning paradigms. Since the training and test datasets consist of instances collected from various do-mains, three different strategies of the usage of training datasets were explored: (1) use all available training datasets and build a unified supervised model for all test datasets; (2) se-lect the most similar training dataset and sep-arately construct a individual model for each test set; (3) adopt multi-task learning frame-work to make full use of available training set-s. Results on the test datasets show that using all datasets as training set achieves the best av-eraged performance and our best system ranks 15 out of 73. 1

Read the paper · More papers on PaperTik