Word Embedding by Combining Resources and Integrating Techniques
Kazem Qazanfari, Abdou S. Youssef · 2019
In a typical text mining problem, the distribution of the domain-specific repository data does not fully capture all the problem's concepts in the real-world. This issue, termed inadequacy of knowledge, decreases the accuracy of generated models. One aspect of inadequacy of knowledge is the out-of-vocabulary problem, i.e. when a word does not appear in the repository data. In this paper, the out-of-vocabulary issue in GloVe is addressed by changing the form of fed training data, i.e. the n-grams of each word are substituted for the word. This version of GloVe is called here C-GloVe. It is shown that the accuracy of the generated models by C-GloVe is mostly higher than when GloVe or FastText is used, especially for smaller training sets. Also, the issue of inadequacy of knowledge is addressed by proposing a method to integrate local (i.e. domain specific) and universal sources of knowledge and to combine different word embedding algorithms. Our experimental results on three different tasks show that the proposed methods yield higher performance than a standalone source of knowledge and a standalone word embedding algorithm, especially if one algorithm of the combination is trained on the local source and another on the universal source of knowledge. Also, experimental results on the classification task show that the proposed method obtained the same or higher F1-score than BERT in four out of five classification problems.