The Acquisition of Khmer-Chinese Parallel Sentence Pairs From Comparable Corpus Based on Manhattan-BiGRU Model
Haiyang Chi, Xin Yan, Siyuan Li, Feng C. Zhou, Guangyi Xu, Lei Zhang · 2020
Bilingual parallel sentence pairs are an important resource for cross-lingual processing. Aiming at the shortage of Khmer-Chinese parallel texts and the complex structure of research methods, a method of obtaining Khmer-Chinese parallel sentence pairs from comparable corpus based on Manhattan-BiGRU model is proposed. In this method, feature information and bilingual word embedding are concatenated together and then jointly used as input. Then the word embedding is coded into sentence embedding through the BiGRU network. Finally, the similarity of sentence embedding is calculated according to the Manhattan distance algorithm, it achieves the acquisition of parallel sentence pairs. Compared with other methods of obtaining parallel sentence pairs based on neural network model, the experimental results show that this method achieves good results, in which the accuracy of obtaining parallel sentence pairs reaches 82.2% and the recall rate reaches 70.7%.