Multimodal Fusion for Coherent Description Generation: A System Integrating NLP, Computer Vision, and Speech Recognition

Priyanka P, T. Balachander · 2025

With the increasing need for intelligent systems capable of understanding and interpreting diverse inputs, multimodal fusion has come out as a promising approach to collaate data from multiple modalities. This paper presents a Multimodal Fusion System for Coherent Description Generation, combining Natural Language Processing (NLP), Computer Vision (CV), Speech Recognition.. This model anchors Convolutional Neural Networks (CNNs) for feature extraction from photos and uses transformer dependant architectures for sequence modeling and coherent text production. The proposed model is structured majorly into three main stages: (1) Multimodal Feature Extraction, (2) Fusion Layer, and (3) Transformer-Based Decoder. Image features are drawn out using CNNs, speech inputs are fabricated using Mel-frequency cepstral coefficients (MFCCs), and sign language data is shown using pose estimation and 3D keypoints. These heterogeneous inputs are lined-up and integrated using a multi-head self-attention mechanism, enhancing cross-model communications. A Transformer decoder then creates a coherent textual description that catches information from all modalities. for better performance, we employ hyperparameter tuning across transformer depth, attention heads, and embedding size. Experimental outcome show that implementing attention-enhanced CNN layers and transformer decoders predominantly improves the model’s strength to create fluent, coherent, and contextually clear descriptions. This system can be enhanced for assistive technologies, human-system interface, and automated content creation, connecting the gap between multimodal understanding and natural language generation.

Read the paper · More papers on PaperTik