Improved multimodal sentiment detection using stressed regions of audio
Harika Abburi, Manish Shrivastava, Suryakanth V. Gangashetty · 2016
Recent advancement of social media has led people to share the product reviews through various modalities such as audio, text and video. In this paper, an improved approach to detect the sentiment of an online spoken reviews based on its multi-modality natures (audio and text) is presented. To extract the sentiment from audio, Mel Frequency Cepstral Coefficients (MFCC) features are extracted at stressed significant regions which are detected based on the strength of excitation. Gaussian Mixture Models (GMM) classifier is employed to develop a sentiment model using these features. From results, it is observed that MFCC features extracted at stressed significance regions perform better than the features extracted from the whole audio input. Further from the transcript of the audio input, textual features are computed by Doc2vec vectors. Support Vector Machine (SVM) classifier is used to develop a sentiment model using these textual features. From experimental results it is observed that combining both the audio and text features results in improvement in the performance for detecting the sentiment of a review.