BERT-Based Implementation of the Serbian Language POS Tagger
Nikola Vukotić, Suzana Stojković · 2025
This paper presents a BERT-based POS tagger for the Serbian language. It compares three transformer models: BERTić—pre-trained on South Slavic corpora—and two Jerteh models (Jerteh-81 and Jerteh-355), both based on RoBERTa, a refined version of BERT. The models were fine-tuned on a simplified, corrected version of the SrWaC corpus using 96 POS tags. Over 20 experiments explored different hyperparameter configurations. Jerteh-355 achieved the best results with 95.88% accuracy and an$\mathbf{F 1}$score of$\mathbf{0. 9 5 9}$, while Jerteh-81 offered a good trade-off for low-resource environments. The findings confirm the effectiveness of transformer-based approaches for Serbian and suggest directions for future work, including full tagset integration and post-processing techniques.