Advanced Machine Learning Frameworks for Multimodal Image and Speech Signal Processing

P Ramesh Naidu, Saurabh Shandilya, C N Srividya, Shantanu Sudhir Gujar, Supriya Devi, Dankan Gowda V · 2024

Multimodal data processing, especially the fusion of image and speech modality, is important for future human computer interface, medical applications and security surveillance. This research proposes the new machine learning approach for the efficient handling of multimodal data in which feature extraction from images and from the speech signal is obtained by employing CNN and LSTM networks, respectively. These features are combined by a hybrid fusion approach so that the resulting algorithm would be both accurate and fast. CIFAR-10 and LibriSpeech were used to assess the framework; the highest accuracy was 92% whilst the conventional methods like MDNN and EFMS could only reach 86%. Moreover, the analysis showed that the proposed method can take a processing time of 0.15 seconds to taxonomy and 0.45 seconds for speech data, thus making it preferable for real time data. These results illustrate the framework’s efficiency in predicting the multimodal input and demonstrate superiority over the prior methods.

Read the paper · More papers on PaperTik