DrugLM: A Unified Framework to Enhance Drug-Target Interaction Predictions by Incorporating Textual Embeddings via Language Models
Tianyi Li, Zhengyu Fang, Xiaoge Zhang, Kaiyu Tang, Huiyuan Chen, Zhimeng Jiang, Tianxiang Zhao, Rong Xu, Feixiong Cheng, Xiao Li, Jing Li · bioRxiv (Cold Spring Harbor Laboratory) · 2025
Abstract Motivation Accurate prediction of drug–target interactions (DTIs) is central to computational drug discovery, offering the potential to reduce experimental costs and accelerate development timelines. While existing deep learning approaches such as Graph Neural Networks and Transformers have shown promise, they often overlook the rich semantic information embedded in textual descriptions of drugs and targets. These descriptions encode critical biomedical knowledge, including mechanisms of action, biological pathways involved, and therapeutic effects of drugs, which can enhance DTI prediction performance. Results We introduce DrugLM , a unified framework that integrates embeddings derived from large language models (LLMs) into DTI-specific model architectures. DrugLM leverages textual descriptions of drugs and targets to generate semantic embeddings using a range of pretrained LLMs. These embeddings can be seamlessly incorporated into existing DTI models. We systematically evaluate multiple LLMs on benchmark DTI datasets and demonstrate strong performance even without fine-tuning. Moreover, supervised parameter-efficient fine-tuning of the LLMs further improves embedding quality, leading to enhanced prediction accuracy. Notably, a simple multilayer perceptron (MLP) using only LLM-derived embeddings surpasses several established DTI methods, underscoring the power of semantic features. Our findings highlight the practical value of integrating LLMs into DTI pipelines and offer a straightforward recipe for improved drug discovery: LLM embeddings of drugs and targets are both effective and easy to use. Availability Our code and dataset are available at https://github.com/ShPhoebus/DrugLM