Tree and word embedding based sentence similarity for evaluation of good answers in intelligent tutoring system
Emil Brajković, Daniel Vasić · 2017
This article presents an approach to examining the similarity of the sentences. In our approach, Euler algorithm was used to generate a series of words based on tree and Sorensen-Dice coefficient was applied to determine the similarity between compared trees. The emphasis is on defining the similarity between the correct and incorrect answers from the Yahoo Question and Answer of the Non-Factual Data Set. Proposed algorithm was used on two types of trees. First is the constituency tree generated by Stanford CoreNLP, and second is custom-made algorithm that produces second type of tree, called knowledge tree which is derived from parse tree. In our comparison, Zhuang-Sasha algorithm was also used. Second approach that was used for sentence comparison uses Word2Vec model for finding word embedding's and calculating sentence average vector, after that cosine distance was applied to determine similarity between two sentences. Results generated with this method were compared with our method in finding sentence similarity based on knowledge tree. Approach described in this paper can be used in evaluation of correct answers which will be used in our implementation of Intelligent Tutoring System.