BongHope: An Annotated Corpus for Bengali Hope Speech Detection
Tanusree Nath, Vivek Kumar Singh, Vedika Gupta · Research Square · 2023
Abstract The rapid growth of social media usage has resulted in huge volume of online content, which includes instances of spread of negativity through the influence of hateful speeches as well as innumerable instances of social media providing supportive and encouraging content, which is often referred to as hope speech. Hope Speech has been recently defined in the context of online content as a textual content that promotes peace and reflects a hopeful attitude. Recently some research has been conducted on automatic detection of hope speech in different languages, including English, Tamil, Malayalam, Spanish and Kannada. However, to the best of our knowledge, there is no work on hope speech research in Bengali language text. Bengali is a language with more than 210 million speakers[2] and a lot of content in various social media platforms. Therefore, it is important to develop appropriate computational methods for automatic detection of text in Bengali as Hope and Non-Hope speech. One possible reason for lack of hope speech research in Bengali language is perhaps the non-availability of suitable dataset/ corpus for the purpose. This work, therefore, attempts to bridge this research gap by presenting a suitably curated and annotated high quality dataset comprising of tweets in Bengali language for the purpose of hope speech research. Several state-of-the-art computational models are applied on the created dataset and the results obtained confirm the suitability of the dataset for hope speech research. [2] https://www.britannica.com/topic/Bengali-language