Task Independent Fine Tuning for Word Embeddings
Xuefeng Yang, Kezhi Mao · IEEE/ACM Transactions on Audio Speech and Language Processing · 2016
Representation learning of words, also known as word embedding technique, is based on the distributional hypothesis that words with similar semantic meanings have similar context. The selection of context window naturally has an influence on word vectors learned. However, it is found that the word vectors are often very sensitive to the defined context window, and unfortunately there is no unified optimal context window for all words. One impact of this issues is that, under a predefined context window, the semantic meanings of some words may not be well represented by the learned vectors. To alleviate the problem and improve word embeddings, we propose a task-independent fine-tuning framework in this paper. The main idea of the task-independent fine tuning is to integrate multiple word embeddings and lexical semantic resources to fine tune a target word embedding. The effectiveness of the proposed framework is tested by tasks of semantic similarity prediction, analogical reasoning, and sentence completion. Experiments results on six word embeddings and eight datasets show that the proposed fine-tuning framework could significantly improve word embeddings.