CROSS: Feedback-Oriented Multi-Modal Dynamic Alignment in Recommendation Systems
Yang Li, Junpeng Du, Chenzhan Wang, Zunlong Liu, Xiaomin Zhu, Chen Lin · ACM Transactions on Recommender Systems · 2025
Aligning the multi-modal content and ID embeddings is crucial in multi-modal recommendation systems. Existing solutions typically adopt a bidirectional alignment paradigm. Our prior work, FETTLE , challenges this paradigm by proposing a one-way directional alignment at the item level, thus reducing the negative impact of low-quality modalities. However, FETTLE leaves two open questions: (1) when is one-way directional alignment optimal, and (2) how to incorporate collaborative signals to enhance alignment? We present CROSS (feedba C k-o R iented multi-m O dal alignment in recommendation S y S tem), a plug-and-play framework that extends FETTLE by introducing three major advancements. First, we introduce Dynamic Item-Level Alignment , which dynamically calibrates the “strength” of each modality via a variance-based compensation mechanism, mitigating the risk of overshadowing weaker modalities in the early stages of training. Second, we develop Multi-grained Collaborative Alignment , which introduces a medium-granularity alignment strategy based on neighboring items that share similar user feedback profiles. This neighbor-level alignment effectively balances noisy user interactions and excessive smoothing across items. Third, we conduct extensive experiments on more real-world datasets and show that CROSS significantly boosts the performance of both collaborative filtering (CF) models and multi-modal recommendation (MRS) approaches, achieving 21.52%–70.78% average improvement on CF backbones and 8.70%–20.73% on MRS backbones. Compared with FETTLE , CROSS achieves additional improvements of 3.82%–5.24%.