An alternative scheme for perplexity estimation
Frédéric Bimbot, Marc El-Bèze, M. Jardino · 2002
Language models are usually evaluated on test texts using the perplexity derived directly from the model likelihood function. In order to use this measure in the framework of a comparative evaluation campaign, we have developed an alternative scheme for perplexity estimation. The method is derived from the Shannon (1951) game and based on a gambling approach on the next word to come in a truncated sentence. We also use entropy bounds proposed by Shannon and based on the rank of the correct answer, in order to estimate a perplexity interval for non-probabilistic language models. The relevance of the approach is assessed on an example.