Multimodal Transformer Fusion for Sentiment Analysis using Audio, Text, and Visual Cues

S. Peerbasha, Mohammed Ihsan Habelalmateen, A C Ramachandra, Revathi P, Thangavelu Saravanan · 2025

In modern era, the increased growth in social media platforms and technologies such as Artificial Intelligence (AI) have gained interest towards multimodal sentiment analysis that includes text, audio and visual cues for extraction of useful insights. Although these systems are capable of analyzing sentiment, but also faces certain challenges in synchronization of multimodal inputs, data fusion and reliability under various contexts. Hence, this research proposes an effective Multimodal Transformer Fusion (MTF) which combines the strengths of various modalities to recognize human emotions. Initially, the data is collected from CMU-Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) dataset. Further, the data is preprocessed with Empirical Mode Decomposition (EMD) and Min-max normalization, stop words removal, and lemmatization to remove noise and improve the quality of sentiment analysis. Then, the features are extracted for different input modalities by using Mel-Frequency Cepstral Coefficients (MFCC), 3D Residual Network (ResNet 3D) and Distilled Bidirectional Encoder Representations from Transformers (DistilBERT). After that, extracted features are further fed into proposed MTF to classify the sentiment of a customer into positive, negative or neutral. From the results, the proposed MTF achieved outstanding results in F1 Score (80.76%) as well as Mean Absolute Error (0.376) when compared with the existing Tensor Fusion Network (TFN).

Read the paper · More papers on PaperTik