Research on Multimodal Data Fusion Model Based on Artificial Intelligence Large Model
Yong Wang, Ting Zhang · 2025
This study explores a multimodal data fusion model based on a large AI model, aiming to enhance information processing capabilities by integrating different types of data such as text, images, and audio. The core algorithms include a pretrained language model (BERT) for text analysis, a convolutional neural network (CNN) for image feature extraction, and a long short-term memory network (LSTM) for processing time-series audio data. In order to optimize the fusion effect between modalities, we applied the Attention Mechanism to dynamically adjust the weights of different data sources to ensure that key information received more attention. In addition, this paper proposes an Adaptive Resource Allocation Strategy (ARAS), which intelligently allocates computing resources according to the quality and importance of each modal data to improve the overall efficiency and accuracy. Specifically, ARAS prioritizes providing more computational resources for more information-rich and reliable modalities through real-time evaluation of input data. Experimental results show that the proposed model performs well in complex scene understanding tasks and achieves a significant surpass over the single-modality method.