Calculation of Chinese-Thai Cross-Language Similarity Based on Sentence Embedding

Feng Yinhan, Zhan Gang, Mao Weixiu, Lin Shunbao, Shijie Yu, Kui Zhang · 2020 5th International Conference on Smart Grid and Electrical Automation (ICSGEA) · 2020

In view of the poor accuracy, efficiency and scalability of the existing cross-language sentence similarity calculations, the Chinese- Thai cross-language sentence similarity is less studied. A new method is proposed. First, preprocess the corpus for Chinese-Thai parallel sentences, the sentence embedding model is used to obtain the Chinese-Thai sentence embedding matrix, and the sentence embedding is normalized. Then, the cross-language mapping model is used to embed the sentence. The conversion matrix is obtained through processing, and the orthogonal optimization of the conversion matrix is performed. Finally, the Chinese sentence embedding is mapped to the Thai sentence embedding space, and the Chinese- Thai cross-language sentence similarity is obtained by calculating the cosine of the two vectors, which provides a new idea for the Chinese-Thai cross-language sentence similarity calculation. The experimental results show that the method has good accuracy.

Read the paper · More papers on PaperTik