Kaz-RoBERTa Conversational Technical Report
Beksultan Sagyndyk, Sanzhar Murzakhmetov, Kirill Yakunin · 2025
We present Kaz-RoBERTa, a base-sized RoBERTa model pretrained from scratch for Kazakh and code-switched Kazakh-Russian conversational text. Motivated by limitations of multilingual models (e.g., mBERT, XLM-R) on real-world, informal language, we curate a 25GB corpus combining a large multi-domain Kazakh dataset and telecom customer-service dialogues. We train a 52k BPE tokenizer and pretrain with masked language modeling on up to 512-token sequences. The resulting model demonstrates improved masked-LM quality and strong gains on downstream tasks including intent classification and code-switch language detection.