Chinese-Khmer Parallel fragments Extraction from Comparable Corpus Based on Dirichlet Process

Shan Ning, Xin Yan, Yu Nuo, Feng C. Zhou, Qing Sheng Xie, Jin Peng Zhang · Procedia Computer Science · 2020

Aiming at the problems existing in the Chinese-Khmer parallel corpus, such as single field, small scale, and poor timeliness, a method of Chinese-Khmer parallel fragment extraction from comparable corpus based on Dirichlet process is proposed. The method firstly obtains the topic distribution of bilingual comparable corpus through bilingual topic model, then uses Poisson distribution to randomly divide the bilingual texts, sets up a threshold to initially filter parallel fragment of comparable corpus, and then obtains the matching probability between parallel fragments by Dirichlet process. Obtaining the final parallel fragments by Gibbs sampling. From the comparison experiments, the method of parallel fragment extraction from bilingual comparable corpus based on Dirichlet process can obtain higher quality parallel fragments without providing any parallel data, which is more suitable for languages where bilingual resources are low.

Read the paper · More papers on PaperTik