Multimodal Fusion for Abusive Speech Detection Using Liquid Neural Networks and Convolution Neural Network
K. S. Paval, Vishnu Radhakrishnan, Keerthana Krishnan, G. Jyothish Lal, B. Premjith · 2024
The recent surge in the use of social media has created vast spaces for viral content that attracts attention of large crowds. This has paved the way to the misuse of these platforms making them breading grounds for toxicity and harassment. Hence there is a need for effective abuse detection methods. In our research, we leverage the ADIMA dataset to investigate abuse detection methodologies, aiming to enhance the effectiveness of existing systems. We propose a multimodal, multilingual abuse detection system that includes three main aspects: the utilization of multimodal fusion techniques for abuse detection, the application of Liquid Neural Networks (LNN) in identifying abusive text content, the use of Convolutional Neural Networks (CNN) in identifying abusive audio utterances and the extension of multimodal fusion abuse detection to cross-lingual settings. This enabled our system to detect abuses in 10 Indian languages. Our approach takes the existing works a step further as it is able to accommodate multiple modalities and multiple languages. We use Convolutional Neural Networks to analyze sound patterns from melspectrograms and Liquid Neural Networks to process text information. To consolidate the strengths of both the representation, late fusion is applied to combine the results, resulting in an ensemble model which achieves an accuracy of 77.47%, and an AUC of 77.89% on the multilingual test set.