Analysis of Word-level Embeddings for Indic Languages on AI4Bharat-IndicNLP Corpora

Dipam Goswami, Shrikant Malviya, Rohit Mishra, Uma Shanker Tiwary · 2021 IEEE 8th Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON) · 2021

This paper presents the analysis of non-contextual word embeddings trained on AI4Bharat-IndicNLP corpus containing 2.7 billion words covering 10 Indian languages. We share the pre-trained embeddings for research and development in Indic languages. These embeddings are evaluated on several evaluation tasks like word similarity and analogy evaluation, classification tasks on multiple datasets. The analysis of word embeddings is expected to give researchers a better understanding of the Indic Languages. We show that Word2Vec skip-gram and FastText skip-gram embeddings are the best performing models for NLP tasks on Indic languages. All the embeddings are made freely available.

Read the paper · More papers on PaperTik