Statistical and Semantic Features to Measure Sentence Similarity in Portuguese

Anderson Pinheiro, Rafael Ferreira Mello, Máverick André Dionísio Ferreira, Vitor Rolim, Joao Vitor S. Tenorio · 2017

A sentence similarity measure is an important field for different applications of text mining. In recent literature, it is possible to find several similarity measures between sentences in English; however, it lacks measures for Portuguese. In addition, one of the main issues to assess sentence similarity is to identify word meaning. In this context, this work aims to present a new approach to measure the similarity between sentences written in Portuguese using statistical and deep learning features to overcome the meaning problems. The results showed that our method obtained better results when compared to the measures proposed in ASSIN 2016 competition.

Read the paper · More papers on PaperTik