Generating Distributed Representations of Words from Japanese Wiktionary with BERT

Ryota Nishiura, Seiji Tsuchiya, Hirokazu Watabe · 2023

In this paper, we introduce a novel approach to create distributed representations of words. While existing methods such as Word2Vec learn from unstructured datasets of plain text, our approach leverages the structured contents of Wiktionary. It employs a pre-trained BERT model to obtain representation vectors of individual Wiktionary headwords by text embedding the corresponding contents of Wiktionary entries. We demonstrate that our proposed method outperforms a Word2Vec model in a XABC test result, potentially due to its ability to leverage both the quality of a pre-trained BERT model and expert-moderated Wiktionary, whereas further research is needed to fully understand the effectiveness of our proposed method under various conditions.

Read the paper · More papers on PaperTik