Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches
Talha Çolakoğlu, Umut Sulubacak, Ahmet Cuneyd Tantug · 2019
With the growth of the social web, usergenerated text data has reached unprecedented sizes.Non-canonical text normalization provides a way to exploit this as a practical source of training data for language processing systems.The state of the art in Turkish text normalization is composed of a tokenlevel pipeline of modules, heavily dependent on external linguistic resources and manuallydefined rules.Instead, we propose a fullyautomated, context-aware machine translation approach with fewer stages of processing.Experiments with various implementations of our approach show that we are able to surpass the current best-performing system by a large margin.