Bayesian Optimization of Text Representations

Dani Yogatama, Lingpeng Kong, Noah A. Smith · 2015

When applying machine learning to problems in NLP, there are many choices to make about how to represent input texts.They can have a big effect on performance, but they are often uninteresting to researchers or practitioners who simply need a module that performs well.We apply sequential model-based optimization over this space of choices and show that it makes standard linear models competitive with more sophisticated, expensive state-ofthe-art methods based on latent variables or neural networks on various topic classification and sentiment analysis problems.Our approach is a first step towards black-box NLP systems that work with raw text and do not require manual tuning.

Read the paper · More papers on PaperTik