Vietnamese- English Cross-Lingual Paraphrase Identification Using Siamese Recurrent Architectures

Le Thanh Nguyen, Điền Đinh · 2019

The paraphrase identification task is one of the important problems that affect the quality of many natural language processing tasks such as querying information, text summarization, plagiarism detection, etc. Especially in the present time, with the development of machine translation tools, the task of paraphrase identification also has to be considered in the case of pairs of texts in two different languages. In this paper, we propose to use the Siamese Recurrent architectures to identify the Vietnamese- English cross-lingual paraphrase cases. In addition, we also use methods such as mapping bilingual word embedding, adding POS vector to word embedding and adjusting the POS tagging label between Vietnamese text and English text. The experimental results in the English- Vietnamese paraphrase corpus with 44,652 sentence pairs show that the use of these methods help to improve the accuracy of the model from 86.86% to 89.61%.

Read the paper · More papers on PaperTik