Hybrid Selection of Language Model Training Data Using Linguistic Information and Perplexity
Antonio Toral · 2013
We explore the selection of training data for language models using perplexity. We introduce three novel models that make use of linguistic information and evaluate them on three different corpora and two languages. In four out of the six scenarios a linguistically motivated method outperforms the purely statistical state-of-theart approach. Finally, a method which combines surface forms and the linguistically motivated methods outperforms the baseline in all the scenarios, selecting data whose perplexity is between 3.49 % and 8.17 % (depending on the corpus and language) lower than that of the baseline. 1