Vowel and Diacritic Restoration for Social Media Texts
Kübra Adalı, Gülşen Eryiğit · 2014
In this paper, we focus on two important problems of social media text normalization, namely: vowel and diacritic restoration.For these two problems, we propose a hybrid model consisting both a discriminative sequence classifier and a language validator in order to select one of the morphologically valid outputs of the first stage.The proposed model is language independent and has no need for manual annotation of the training data.We measured the performance both on synthetic data specifically produced for these two problems and on real social media data.Our model (with 97.06% on synthetic data) improves the state of the art results for diacritization of Turkish by 3.65 percentage points on ambiguous cases and for the vowel restoration by 45.77 percentage points over a rule based baseline with 62.66% accuracy.The results on real data are 95.43% and 69.56% accordingly.