Combining conditional random fields and word embeddings to improve Amazigh part-of-speech Tagging
Rkia Bani, Samir Amri, Lahbib Zenkouar, Zouhair Guennoun · 2023
Part-of-speech tagging is considered a key task in natural language processing projects. Part-of-speech taggers that use classical conditional random fields perform very well in well-studied and rich language reaching 97%. However, in low-resourced language like Amazigh language, the accuracy of the conditional random fields model with features extraction is less than 90% in tagging. In this paper, we present a new part-of-speech tagging model using conditional random fields with the integration of word-embeddings. This method uses word-embeddings as features instead of the function of extracting features that is usually used in conditional random fields models. Because the Amazigh language suffers from the absence of pretrained word-embeddings, we designed a word-embedding layer that is trained gradually as the model. We used an existing dataset of 60k tokens and a tagset of 28 tags to evaluate our model. The accuracy of our proposed model reaches 96.78 %. The result is compared with other state-of-art Amazigh part-of-speech taggers that used conditional random fields. This comparison shows that our approach outperforms the other techniques in terms of accuracy. In the future, will continue enriching the Amazigh language with more natural language processing applications and datasets.