An evaluation of statistical language models of spoken dialogue using the British National Corpus
G. J. A. Hunter · 2005
The use of spoken dialogue interfaces is becoming quite commonplace in many everyday situations such as cinema ticket bookings and much effort is being put into making such systems reliable and easy to use. There is strong evidence that the performance of the language model component of a speech recognition system is heavily dependent on the nature of the material to which it is applied relative to the nature of the material on which it was trained. Although much work has been carried out on the statistical modelling of text data based on, for example, transcripts of news broadcasts, relatively little work has been carried out to date on applying such models to dialogue data. This would seem to be an important area for study with a view improving automated spoken dialogue interfaces. This paper describes work applying statistical modeling techniques to both text and dialogue material in the British National Corpus, comparing the results obtained from the two distinct datasets and interpreting these findings in the light of psycholinguistic theories of dialogue.