MultiSMDTA: Drug–Target Affinity Prediction Based on Multi-Scale Multimodal Features and Multi-Head Self-Attention
Jiayan Lu · 2025
Drug–target affinity (DTA) prediction plays a crucial role in drug discovery advancement. In recent years, deep learning methods have become widely used in DTA tasks. Drugs and targets can be represented through various approaches, including structure-based, sequence-based, and graph-based representations. Nevertheless, most existing studies only use single-modality or single-scale representations for drug molecules or protein targets. This leads to limited feature information. Even when attempting to integrate multimodal information, current methods typically apply simple strategies like feature concatenation or weighted summation. Such approaches do not fully capture the complex interactions and nonlinear relationships between modalities. Therefore, multimodal information integration remains insufficient, restricting improvements in model performance. To address these challenges, we propose a novel multi-scale and multimodal drug–target affinity prediction model named MultiSMDTA. By incorporating multi-scale and multimodal information of proteins, MultiSMDTA effectively captures both the biological sequence and structural information of protein targets. Furthermore, a multimodal fusion module based on multi-head self-attention is designed to better model interactions among different modalities. Experimental results on the Davis and KIBA benchmark datasets demonstrate that MultiSMDTA consistently outperforms existing state-of-the-art methods across multiple evaluation metrics.