Batch IS NOT Heavy: Learning Word Representations From All Samples

Xin Xin, Fajie Yuan, Xiangnan He, Joemon M. Jose · 2018

Stochastic Gradient Descent (SGD) with negative sampling is the most prevalent approach to learn word representations.However, it is known that sampling methods are biased especially when the sampling distribution deviates from the true data distribution.Besides, SGD suffers from dramatic fluctuation due to the onesample learning scheme.In this work, we propose AllVec that uses batch gradient learning to generate word representations from all training samples.Remarkably, the time complexity of AllVec remains at the same level as SGD, being determined by the number of positive samples rather than all samples.We evaluate AllVec on several benchmark tasks.Experiments show that AllVec outperforms samplingbased SGD methods with comparable efficiency, especially for small training corpora.

Read the paper · More papers on PaperTik