Enhancing Tokenization by Embedding Romanian Language Specific Morphology
Mihaela Alexandra Vasiu, Rodica Potolea · 2020
this paper addresses the significance, complexity and approaches of Romanian tokenization, the first step in NLP. A strategy is proposed for tokenizing the Romanian text with various specificity: standard (literal language), archaic and medical. The strategy is extending existing tokenization models for English, Spanish and Portuguese with Romanian features. Moreover, we attempt to address the Romanian ambiguities caused by the complexity of the grammar and provide support for further Romanian NLP processing. The approach proved that there is no general solution for tokenization and that diverging tokenization strategies for all languages can lead to considerably better results. With the current approach we obtained a high-quality tokenization on all three Romanian formats as follows:for standard sources,for archaic sources andfor medical sources.