A Lightweight Multimodal Learning Model to Recognize User Sentiment in Mobile Devices

Jyotirmoy Karjee, N. N. Srinidhi, Gargi Dwivedi, Arun Padmanabh Bhagavath, Prajwal Ranjan · 2023 IEEE International Conference on Consumer Electronics (ICCE) · 2023

Communications through video/audio calling and text messaging has seen a meteoric rise over the last few years. As we progress towards the advanced wireless technology (i.e., 5G), the quality of call and network low latency communication provides better quality of services (QoS). This allows for addition of new features that would provide users with a more personalized experience such as providing sentiments analysis of users through mobile (sender/receiver) devices over text messaging or video calls or both. However to perform sentiment analysis requires Deep Neural Networks (DNN) model to extract multimodal data (such as text, audio and video) features which is very heavy to deploy in mobile devices due to its limited computational capabilities. In the past, various research works has been performed to develop multimodal feature extraction mechanisms, however none of the multimodal models are suited to be deployed in mobile devices for practical applications. To mitigate these issues, we propose a light weight multimodal learning model called Tri-Feature Fusion which can be easily deployable in mobile devices for client-server communications. The Tri-Feature Fusion model extracts the feature vectors of each modality and generate a single multimodal feature vector from the data generated from the call and perform sentiment analysis on each particular mode. We also have developed a light weight neural network for Tri-Feature Fusion to perform the sentiment analysis using the extracted features for real-time audio, video and text data displayed at clients and receiver mobile devices. We conduct extensive experiments to showcase the performance of Tri-Feature Fusion as light weight model in-terms of average time taken for prediction, time taken per step, CPU utilization and size of the models compared with the existing state of art.

Read the paper · More papers on PaperTik