Towards Amazigh Word Embedding: Corpus Creation and Word2Vec Models Evaluations
Hassan Faouzi, Maria El-Badaoui, Mohammed Boutalline, Adil Tannouche, Hamid Ouanan · Revue d intelligence artificielle · 2023
Distributed representations of words in a vector space help learning algorithms to model semantic notions of word similarity and distances in sentences.Most of the existing researches have been done on the Latin, Arabic, and other language, while the Amazigh language is ignored.In this paper, we try to build a first model word embeddings for Amazigh language and describe the steps needed to build it.Therefore, we implement a Word2Vec that a combination of two techniques -CBOW (Continuous bag of words) and Skip-gram to transform words written in Tifinagh to vector form.To obtain the highest performance, we evaluate two parameters of Word2Vec include Word2Vec model architecture and vector dimension.This evaluation process was implemented towards our proposed corpus collected on Amazigh websites for different domains.The result shows that the highest accuracy values are obtained under the combination of CBOW model and 300 dimensional vector.