A Cascaded Approach for Social Media Text Normalization of Turkish
Dilara Torunoğlu, Gülşen Eryiğit · 2014
Text normalization is an indispensable stage for natural language processing of social media data with available NLP tools.We divide the normalization problem into 7 categories, namely; letter case transformation, replacement rules & lexicon lookup, proper noun detection, deasciification, vowel restoration, accent normalization and spelling correction.We propose a cascaded approach where each ill formed word passes from these 7 modules and is investigated for possible transformations.This paper presents the first results for the normalization of Turkish and tries to shed light on the different challenges in this area.We report a 40 percentage points improvement over a lexicon lookup baseline and nearly 50 percentage points over available spelling correctors.