Incorporating Molecular Knowledge in Large Language Models via Multimodal Modeling
Zekun Yang, Kun Lv, Jian Hong Shu, Zheng Li, Ping Xiao · IEEE Transactions on Computational Social Systems · 2025
In recent years, large language models (LLMs) represented by GPT-4 have achieved tremendous success in natural language-centered tasks. Nevertheless, LLMs face inherent challenges in tasks involving both natural language and molecular modalities. Although there has been some research progress on these tasks, two challenges remain unresolved: modeling the differences in representation format between natural language and molecular modalities, and capturing subtle differences under an instruction-tuned paradigm. To address these two challenges, this article proposes a two-stage training framework to build molecular knowledge-enhanced LLM, named Mol-LLM. The first stage utilizes the multitask instruction tuning method to tackle the modality differences between natural language and molecular sequence. The second stage is the direct preference optimization training strategy with three random preference actions to capture subtle differences under the instruction-tuned paradigm. Extensive experiments have demonstrated the state-of-the-art performances of the model Mol-LLM proposed in this study, includingmol2mol,text2text,mol2text, andtext2moltask. The effectiveness of the module has been further verified through ablation studies, and the generalizability has been confirmed by additional supportive experiments.