Robust Learning in Random Subspaces: Equipping NLP for OOV Effects

Anders SÃ ̧gaard, Anders Johannsen · International Conference on Computational Linguistics · 2012

Inspired by work on robust optimization we introduce a subspace method for learning linear classifiers for natural language processing that are robust to out-of-vocabulary effects. The method is applicable in live-stream settings where new instances may be sampled from different and possibly also previously unseen domains. In text classification and part-of-speech (POS) tagging, robust perceptrons and robust stochastic gradient descent (SGD) with hinge loss achieve average error reductions of up to 18% when evaluated on out-of-domain data.

Read the paper · More papers on PaperTik