A Survey of Using Unsupervised Learning Techniques in Building Masked Language Models for Low Resource Languages
Labehat Kryeziu, Visar Shehu · 2022 11th Mediterranean Conference on Embedded Computing (MECO) · 2022
A very common approach nowadays towards learning how to represent text in machines is to use transformers. These models are based on neural networks, and they show promising results when applied to problems in which both the input and output of the model are sequences. In this paper we give an overview on what transformers are with a focus on Bidirectional Encoder Representations from Transformers (BERT). Furthermore, we analyze different approaches on how these models are used for text representation in low resource languages. The final goal is to establish whether BERT can be applied to establish NLP capabilities for the Albanian language.