Word prediction techniques for user adaptation and sparse data mitigation
Kathleen F. McCoy, Keith Trnka · 2011
The field of Augmentative and Alternative Communication (AAC) seeks to help minimize the effects of speech disorders to allow better communication. High-tech AAC devices are electronic devices that address the dual issue of speech impairment and reduced motor control. Communication rate can be improved substantially using word prediction, an application of language modeling to text entry. Word prediction relies on the letters and words that the user has entered so far to suggest likely words that the user is in the process of typing. Although we can increase language model quality with more training data, these efforts are unlikely to increase the average relevance of training texts. Instead, many documents will be dissimilar to testing data. Even if a few relevant texts are added to the training data, their contribution is marginalized by the abundance of irrelevant texts. We address the problem of varying relevance through adaptive language modeling, specifically topic adaptation and style adaptation. We found that these adaptations increase keystroke savings for both topic and style adaptation individually, and also when topic and style modeling are combined. We have addressed the problem of irrelevant training data with cache models to learn new words and model lexical repetition, and we also integrated a large word list using methods developed for the part of speech ngram model. We have focused on general-purpose improvements in language modeling and natural language processing. Many of our methods may be applicable to related language modeling problems. However, we have focused only on language modeling improvements which we feel are well-suited for word prediction and AAC.