Sentiment Analysis of Sindhi News Articles using Deep Learning

Fahama Barakzai, Sania Bhatti, Salahuddin Saddar · 2022 IEEE 17th International Conference on Computer Sciences and Information Technologies (CSIT) · 2022

Sindhi is an Indo-Aryan language that is spoken by 24 Million people around the globe. By analyzing the recently researched data one can find that this is the lowest researched language with regard to the computations of the digital era, lacking behind in most importantly the sentiment analysis technique of Natural Language Processing. The digitization of the world required this language to be analyzed for prediction and other purposes. Sindhi can be analyzed thoroughly through Sentiment Analysis and several Natural Language Processing models are contributing to making the languages more predictable. The insubstantial amount of text analysis of this language can be because of the complexity of this language. Hence, there is a need for this language to be analyzed more efficiently using extraordinary machine learning algorithms. This paper focuses on the major aims of analysis of the Sindhi language’s Newspaper Dataset ‘Awami Awaz’ using the very famous machine learning algorithm Bi-Directional Encoder Representation from Transformers (BERT). This model works in both directions and accepts 104 languages. This work also claims some challenges in analyzing the typical Sindhi language which is not a chunk of BERT’s multilingual 104 languages dataset. The problem is tackled by transliterating this language into English for achieving results. BERT, after translation through Google Translation API, secured 67.2% of accuracy. The hindrances can be overcome by adding the Sindhi vocabulary to BERT’s multilingual dataset.

Read the paper · More papers on PaperTik