Sentiment Analysis of Custom Speech Corpus: A proof of concept for NLP

Suja Sreejith Panickar, Rimjhim Sinha, Vidhi Chawla, Omkar Singh, Omkar Londhe · Procedia Computer Science · 2024

This paper presents an innovative, custom Hindi Speech Corpus dedicated to assessing the positive intensity in Hindi language. By meticulously gathering voice samples from individuals across diverse demographics, this corpus is well crafted for acoustic research in Sentiment Analysis. Several sentences spanning a spectrum of sentiments (extremely positive, positive, neutral, negative) were acquired from participants. The resulting corpus comprises of 225 audio files. To guarantee the consistency and to accommodate domain knowledge based nuances, the sentiments and its intensity (high, medium, low) were classified by Expert annotation. We used PowerBI for attaining insights into emotional landscape of the corpus. Employing exploratory analysis and Sentiment Analysis, the corpus is analyzed to unveil linguistic features indicative of positive intensity in Hindi speech. This study advances the understanding of positivity expression in Hindi and furnishes valuable resources for linguistic investigations in the scope of Hindi language. It is observed that positive sentiment has the highest count at 58 and the lowest count is for very negative at 27. The main contribution of current work lies in creating the custom dataset in Hindi which to our knowledge is first of its kind. Also, we did not come across any dataset focused on intensity of positivity. These contributions shall definitely be beneficial to future researchers in exploring positive affect related research.

Read the paper · More papers on PaperTik