Tibetan-Chinese cross-lingual word embeddings based on MUSE

Wei Ma, Hongzhi Yu, Kun Zhao, Deshun Zhao, Jun Yang · Journal of Physics Conference Series · 2020

Abstract The idea of word embedding is based on the semantic distribution hypothesis of linguist Harris (1954), who believes that words with the same semantics are distributed in similar contexts. The learning of word embedding is a crucial technology in natural language processing. In recent years, cross-language word vectors have received more and more attention. Cross-language Word vectors can transfer knowledge between different languages. Most importantly, this transfer can occur between rich-resource and low-resource languages. This paper uses the Tibetan-Chinese Wikipedia corpus to train monolingual word vectors. Based on the Tibetan-Chinese bilingual translations, we use the supervised method in the MUSE library to train the Tibetan-Chinese bilingual cross-language word embeddings. In the experiment, we evaluate the result of word representation on the standard lexical semantic evaluation task. The results show that the method has a certain improvement in the semantic representation of the word embedding.

Read the paper · More papers on PaperTik