MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response

Zihao Deng, Yinghao Ma, Yudong Liu, Rongchen Guo, Ge Zhang, Wenhu Chen, Wenhao Huang, Emmanouil Benetos · 2024

Large Language Models (LLMs) have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored.To address this gap, we present MusiLingo, a novel system for music caption generation and musicrelated query responses.MusiLingo employs a single projection layer to align music representations from the pre-trained frozen music audio model MERT (Li et al., 2023b) with a frozen LLM, bridging the gap between music audio and textual contexts.We train it on an extensive music caption dataset and fine-tune it with instructional data.Due to the scarcity of highquality music Q&A datasets, we created the MusicInstruct (MI) dataset from captions in the MusicCaps datasets, tailored for open-ended music inquiries.Empirical evaluations demonstrate its competitive performance in generating music captions and composing music-related Q&A pairs.Our introduced dataset enables notable advancements beyond previous ones.

Read the paper · More papers on PaperTik