MARS: Multimodal-Assisted Refined Semantic Alignment

Junjie Xu, Xingjiao Wu, Zihao Zhang, Shuwen Yang, Tianlong Ma, Daoguo Dong, Liang He · Information Processing & Management · 2025

Audio-to-image generation (AIG) faces challenges in fine-grained semantic alignment, particularly semantic semantic misalignment, and loss of visual detail. To address these issues, we proposed MARS ( M ultimodal- A ssisted R efined S emantic alignment), a novel framework leveraging a Mamba-based audio encoder to manage the complexity of long audio sequences, coupled with a fine-grained multimodal alignment strategy using visual descriptions from multimodal large language models. We enhanced semantic coherence and aesthetic quality by fine-tuning an image generator using an image aesthetic perception generator. Furthermore, we validated MARS on VGGSound and VEGAS benchmarks, comprising 37,250 and 9,500 records, respectively. The results suggest that MARS significantly outperforms existing methods, achieving average improvements of 28.73% in semantic relevance and 127.35% in aesthetic scores compared with the best AIG generation baseline. In addition, cross-domain evaluations on the AudioCaps and Clotho datasets confirmed the robustness and generalization capability of MARS , with an average improvement of 73.9% on the V2A metric. • MARS refines semantic alignment for audio-based image generation. • A Mamba-based encoder processes long audio sequences. • MLLMs supply rich visuals to boost semantics and aesthetics. • Extensive experiments prove effectiveness and robustness.

Read the paper · More papers on PaperTik