Transformer-Based Language Models for Bulgarian

Iva Marinova, Kiril Simov, Petya Osenova · 2023

This paper presents an approach for training lightweight and robust language models for Bulgarian that mitigate gender, political, racial, and other biases in the data.Our method involves scraping content from major Bulgarian online media providers using a specialized procedure for source filtering, topic selection, and lexicon-based removal of inappropriate language during the pre-training phase.We continuously improve the models by incorporating new data from various domains, including social media, books, scientific literature, and linguistically modified corpora.Our motivation is to provide a solution that is sufficient for all natural language processing tasks in Bulgarian, and to address the lack of existing procedures for guaranteeing the robustness of such models.We evaluated the performance of our language models on several Natural language processing (NLP) tasks, including filling the mask, text generation and named entity recognition (NER).We also performed bias analysis on our models to ensure that they are not biased towards any particular group or ideology.Our analysis showed that within our setting the models have a low level of bias towards gender, race, etc. Needless to say, more experiments have to be performed in future that incorporate comparison with non-biased data and relies on more bias-related prompts.

Read the paper · More papers on PaperTik