Chinese-Vietnamese cross-language topic discovery method based on generative adversarial networks
Xia Linjie, Zhengtao Yu, Gao Shengxiang · 2022
The cross-language news topic discovery task aims to cluster news texts in different languages that describe the same topic and classify the topic in the form of keywords. At present, most cross-language topic discovery methods are based on machine translation or external resources like bilingual dictionaries and parallel sentences to solve cross-language problems. However, Vietnamese is a low resource language and it is difficult and expensive to manually annotate ChineseVietnamese bilingual aligned corpora. To solve this problem, this paper proposes a Chinese-Vietnamese cross-language topic discovery method based on generative adversarial networks (GAN). Firstly, News texts are represented as vectors by BERT, and then the bilingual vectors are mapped to the same semantic space by GAN. Finally, k-means clustering algorithm is used to cluster the representation vectors and extract the topics. Experiments on the Chinese-Vietnamese bilingual news topic discovery corpus show that the proposed method is superior to the baseline.