Mimicking Word Embeddings using Subword RNNs
Yuval Pinter, Robert Guthrie, Jacob Eisenstein · 2017
Word embeddings improve generalization over lexical features by placing each word in a lower-dimensional space, using distributional information obtained from unlabeled data.However, the effectiveness of word embeddings for downstream NLP tasks is limited by out-of-vocabulary (OOV) words, for which embeddings do not exist.In this paper, we present MIM-ICK, an approach to generating OOV word embeddings compositionally, by learning a function from spellings to distributional embeddings.Unlike prior work, MIMICK does not require re-training on the original word embedding corpus; instead, learning is performed at the type level.Intrinsic and extrinsic evaluations demonstrate the power of this simple approach.On 23 languages, MIMICK improves performance over a word-based baseline for tagging part-of-speech and morphosyntactic attributes.It is competitive with (and complementary to) a supervised characterbased model in low-resource settings.