Hierarchical Fusion Framework for Multimodal Dialogue Response Generation
Qi Deng, Lijun Wu, Kaile Su, Wei Wu, Zhiyuan Li, Weiwei Duan · 2024
Two analogous tasks have emerged in multimodal dialogue research: multimodal dialogue response generation and multimodal task-oriented dialogue. Both tasks share the goal of response multi-round, interactive content based on the multimodal dialogue history, but the latter focuses on accomplishing specific objectives which can be viewed as the former fine-tuned. The fine-tuning strategy may cause catastrophic forgetting and overfitting on few well-annotated data. Despite considerable progress in both areas, many existing works rely on retrieval-based approaches and additional auxiliary knowledge bases. To address these issues, we propose a Hierarchical Fusion Framework (HFF) for multimodal dialogue response generation. HFF blends these two tasks to learn a generation model from a data-driven perspective by introducing multi-dataset learning scheme, achieving a balance between generalization and expertise. In this work, multi-dataset learning is cast as a multi-objective optimization problem due to potential conflicts between datasets, necessitating a trade-off based on data distribution during training. Hierarchical fusion is performed sequentially between modalities and datasets, which could efficiently establish clear cross-modal relationships and integrate knowledge from multi-dataset. Specifically, HFF aligns extracted unimodal features (image and text) before fusing them through cross-modal attention and integrates them into multimodal encoder-decoder for generating responses. By optimizing the fusion between corpora from multi-dataset as conflicting objectives to satisfy Pareto optimality, our approach effectively facilitates both multimodal task-oriented and task-unoriented dialogues. Experimental results demonstrate the effectiveness of HFF and its comparable performance with all baselines.