Understanding sarcasm in speech using mel-frequency cepstral coefficent
Abhinav Mathur, Vikas Saxena, Sandeep Kumar Singh · 2017
Multimedia data has been playing an important role in comprehending meaningful information to make it more entertaining and interactive. Most of the video contents may contain sarcastic, ironic or metaphorical conversation. To detect the presence of sarcastic words, words in subtitles of the video are being used to infer the information that is being conveyed. However, downloading subtitles of any video containing sarcasm and the η understanding it becomes a monotonous process. Many times, words in subtitles may not be sarcastic but are spoken sarcastically by changing the tone, pitch etc which is difficult to access by linguistic analysis alone. In this paper, we resolve the this issue of deleting sarcastic speech using a combination of audio mining technique and spectral features. The proposed work has two distinct modules. The first one is Audio Extraction which converts an input audio file of any format to .wav format, followed by second module Sarcastic Speech Recognition of the extracted .wav file using pitch as a major parameter. Our proposed approach has reported accuracy of 70%.