Overfitting at SemEval-2016 Task 3: Detecting Semantically Similar Questions in Community Question Answering Forums with Word Embeddings

Hujie Wang, Pascal Poupart · 2016

This paper presents an approach for estimating the question-question similarity of an English dataset specified in Shared Task 3, subtask B of SemEval-2016. Given a new question and a set of the first 10 related questions retrieved by a search engine, participants are asked to produce a binary relevant/irrelevant judgement and rerank the related questions according to their similarity with respect to the original question. Our submitted system uses a 2-layer feed-forward neural network with the averages of word embedding vectors to predict the semantic similarity score of two questions. We also evaluate the results of Random Forests and Support Vector Machine in comparison to the Neural Network. Results on the test dataset show that the model achieves a Mean Average Precision 69.681.

Read the paper · More papers on PaperTik