NepaliBERT: Pre-training of Masked Language Model in Nepali Corpus
Shushanta Pudasaini, Subarna Shakya, Aakash Tamang, Sajjan Adhikari, Sunil Thapa, Sagar Lamichhane · 2023
Recent improvement in the Transformer model has provided us with state-of-the-art architectures like BERT, RoBERTA, and GPT-3. Though, better results are obtained for the real-world NLP problems with the use of these architectures, there is still a lot to achieve in the Nepali domain. Nepali Language, which uses the Devanagari script, has rich semantics and grammatical structure but due to the lack of computational resources, the optimum results are yet to be achieved by using the state-of-the-art architectures . This is why, there is no publicly available NLP Nepali model, which can be used by other researchers. This research study intends to fill this research gap by providing word embeddings for the Nepali language trained on Word2Vec, Doc2Vec and BERT architecture, which can be used as a base for creating benchmark results on different NLP tasks.