Normalization of text messages for text-to-speech
Deana L. Pennell, Yang Liu · 2010
This paper describes a normalization system for text messages to allow them to be read by a TTS engine. To address the large number of texting abbreviations, we use a statistical classifier to learn when to delete a character. The features we use are based on character context, function, and position in the word and containing syllable. To ensure that our system is robust to different abbreviations for a word, we generate multiple abbreviation hypotheses for each word based on the classifier's prediction. We then reverse the mappings to enable prediction of English words from the abbreviations. Our results show that this approach is feasible and warrants further exploration.