MLFF-DTA: A Multi-Level Feature Fusion Method for Predicting Drug-Target Binding Affinity

Jiao Wang, Ge Kong, Juan Wang · 2024

Accurate prediction of drug-target binding affinity (DTA) is crucial for drug discovery. In deep learning-based methods, robust feature representations of drugs and targets and their interaction features play a key role in improving the accuracy of DTA prediction. Additionally, biological data on DTA have been significantly updated in recent years, and the ability to predict and identify new data is also an important consideration. In this paper, we have constructed two new datasets to update the newly discovered DTA data. We propose a new method based on Multi-Level Feature Fusion for predicting DTA, called MLFF-DTA. It uses SMILES strings of compounds and amino acid sequences of proteins as inputs and then extracts molecular graph and fingerprint features for compounds and n-gram features for proteins. Here, three kinds of complementary fingerprints, i.e., MACCS, PubChem, and Pharmacophore ErG fingerprints, are fused as a new fingerprint feature, called FP, to describe the physicochemical properties of compounds. The n-gram features describe the types and quantities of amino acid functional groups in protein sequences. MLFF-DTA is trained and tested on three datasets, i.e., a public dataset and two new datasets. MLFF-DTA was also tested for the ability of generalization on four small datasets, in which each protein is not included in training sets. The experimental results demonstrate that the performance of MLFF-DTA is superior to other baseline methods. Furthermore, we apply MLFF-DTA to discover three potential drugs for breast cancer with support from relevant clinical experiments. The source code and datasets of MLFF-DTA are available at https://anonymous.4open.science/r/MLFF-DTA.

Read the paper · More papers on PaperTik