Building a Framework for Identifying Arabic Dialects Using Deep Learning Techniques

Mohammed Qassim Shatnawi, Muneer Bani Yassein, Aseel Abu Huq · ACM Transactions on Asian and Low-Resource Language Information Processing · 2023

Statista statistics show that social media usage has rapidly increased over the past ten years, and that by 2021, there will be approximately 4 billion 871 million Internet users worldwide. This is because mobile phone platforms allow users to express their sentiments and opinions in a variety of languages, including Arabic. One of the most widely used social media sites, Twitter offers a conducive environment for users to voice their thoughts. It is used by magazines and government websites to publish official statements and decisions, and while they write in Modern Standard Arabic, consumers interact with it in their native languages. To tackle issues with natural language processing, such as sentiment analysis and translation, a huge volume of data in many languages must be interpreted and processed. We combine a dataset from Katherine that covers the dialects of 8 countries with a dataset from the NADI shared tasks that was gathered via Twitter and includes the dialects of 21 countries. Three deep learning models with various word embeddings were used. The best results were obtained by CNN/BiLSTM with FastText (51.4

Read the paper · More papers on PaperTik