The Effect of Model Capacity and Script Diversity on Subword Tokenization for Sorani Kurdish
Ali Salehi, Cassandra L. Jacobs · 2024
Tokenization and morphological segmentation continue to pose challenges for text processing and studies of human language.Here, we focus on written Soranî Kurdish, which uses a modified script based on Persian and Arabic, and its transliterations into the Kurdish Latin script.Importantly, Perso-Arabic and Latin-based writing systems demonstrate different statistical and structural properties, which may have significant effects on subword vocabulary learning.This has major consequences for frequency-or probability-based models of morphological induction.We explore the possibility that jointly training subword vocabularies using a source script along with its transliteration would improve morphological segmentation, subword tokenization, and whether gains are observed for one system over others.We find that joint training has a similar effect to increasing vocabulary size, while keeping subwords shorter in length, which produces higherquality subwords that map onto morphemes.