Clustering words for statistical language models based on contextual word similarity
Azarshid Farhat, Joanna Isabelle, Douglas D. O’Shaughnessy · 2002
This paper describes a new word clustering approach for statistical language modeling. The classification criteria used by our approach is the contextual word similarity used in a simplified clustering algorithm. This clustering technique was tested on the INRS speech recognizer using the spontaneous English corpora, ATIS. Automatic word classification increases the word accuracy rate by 8.6% with a perplexity reduction about of 6.9%.