Hybrid Selection of Language Model Training Data Using Linguistic Information and Perplexity

Antonio Toral · 2013

We explore the selection of training data for language models using perplexity. We introduce three novel models that make use of linguistic information and evaluate them on three different corpora and two languages. In four out of the six scenarios a linguistically motivated method outperforms the purely statistical state-of-theart approach. Finally, a method which combines surface forms and the linguistically motivated methods outperforms the baseline in all the scenarios, selecting data whose perplexity is between 3.49 % and 8.17 % (depending on the corpus and language) lower than that of the baseline. 1

Read the paper · More papers on PaperTik