English- Vietnamese Cross-Language Paraphrase Identification Method

Le Thanh Nguyen, Điền Đinh · 2017

Paraphrase identification is a very important problem and is used in many natural language processing tasks such as machine translation, bilingual information retrieval, plagiarism detection, etc. With the development of information technology and the internet, the requirement of textual comparing is not only in the same language but also in many different language pairs. Especially in Vietnamese, the need to detect paraphrase in English-Vietnamese pair of sentences is very large because English is a most popular foreign language in Vietnam. However, the in-depth studies on cross-language paraphrase identification task between English and Vietnamese are still limited. In this paper, we propose a method to identify the English-Vietnamese cross-language paraphrase cases using a fuzzy-based method and the BabelNet semantic network. We identify if the pair of sentences is paraphrased by using feature classes and then combine these results into a final one using a mathematical formula. The experimental results show that our model achieves 77.1% F-measure accuracy and has the advantage of the processing speed compared to other methods which have equivalent quality.

Read the paper · More papers on PaperTik