A novel weakly supervised approach for casual English normalization

Assia Mezhar, Mohammed Ramdani, Amal El Mzabi · 2016

Nowadays, social media has become a massive real-time source of fresh information. Which gather a large volume of data produced every day to analysts who wish to discover new opportunities in the emerging research area of mining the noisy corpus present in social media. However, the noisy and short nature of this kind of text hamper this desire: unlike structured news content, social media users often prefer communicating unconventionally with informal, and ungrammatical language using abbreviations, slang, misspelled words, or non-standard short-forms: noisy words. Under those purposes, it becomes a challenge to present new methods to boost the performance of existing systems which convert this noisy text to Standard English. Most previous work didn't encompass all kinds of casual English noise, or made a strong assumption that the best canonical candidate is the one present in the correction dictionaries. This is not realistic because one noisy word can have different meanings considering the context, the area of interest and the time period of those noisy words extraction. In this paper, we target all kinds of casual English noise neglected by state-of-art by proposing a novel weakly supervised approach. This one takes into account the context, the type of interest of those noisy words and their time period of extraction: Once the informal word is identified, we recognize its type of interest by surrounding its context. Then, we generate the best and highly ranked canonical form depending on the extraction time period of the noisy word. Finally, we normalize the noisy word. This multi-faceted, context aware, and interest sensitive approach is expected to yield more accurate results than existing state-of-art solutions.

Read the paper · More papers on PaperTik