Efficient Multimodal Visual Segmentation Model Based on Phased Fusion of Differential Modalities

International Journal of Big Data Intelligent Technology · 2025

This article focuses on the application of an efficient multimodal visual segmentation model based on phased fusion of differential modalities in image harmonization tasks.In response to the problem of the failure of the harmonization model due to the lack of a predetermined foreground mask in practical application scenarios, this paper innovatively proposes a multimodal image harmonization task, which replaces the foreground mask by introducing a referential description for the foreground.This article constructs a new multimodal harmonization dataset ReiHarmony4 based on the traditional image harmonization dataset iHarmony4.For this task, this article proposes a segmentation harmonization pipeline model and two different end-to-end methods, including a combination of CLIP based referential image segmentation model and Harmony Transformer, and DiffHarmony based on stable diffusion model.The experimental results show that these models can effectively complete the task of multimodal image harmonization.This article also designs and implements a multimodal image harmonization system, which achieves the function of image harmonization based on text.Although some achievements have been made, there are still some issues that need further exploration, such as improving segmentation performance, solving the problem of resolution degradation, and expanding the dataset.

Read the paper · More papers on PaperTik