Attention-Based Sequence Learning Model for Arabic Diacritic Restoration

Ali Alkhathlan, Faris Kateb, Jugal Kumar Kalita · 2020

Diacritic restoration is a fundamental issue in many languages such as Arabic, Greek and others. The bulk of online texts lack these markings. This makes these languages difficult for a human to interpret and computer process. In this work, we will automatically restore the diacritics for the Arabic language. Previous sequence-to-sequence based approaches such as Abandah et al., [1], and Belinkov and Glass [2] have been used. These use Recurrent Neural Network (RNN) involving Long- Short Term Memory (LSTM) models. Interestingly, these are language-independent, avoiding any morphological, syntactical analysis or any type of external resources. This approach leads to making this model agnostic to new languages. In this work, we extend the previous work with an alternate RNN architecture involving Gated-Recurrent Units (GRU). We augment the basic RNN model with attention-based mechanisms. Our GRU-attentive model improved the Diacritic Error Rate (DER) by 0.19% and by 2.32% compared to the state of the art.

Read the paper · More papers on PaperTik