Predicting Drug Functions from Gene Ontology, Amino Acid Sequences, and Drug-Disease Associations through Multi-label Machine Learning with MLSMOTE

Pranab Das, Dilwar Hussain Mazumder · 2023

Drug function prediction plays a key role in drug design, discovery, and development, with millions of dollars being spent annually on this time-consuming and costly process. Machine learning-based drug function prediction has emerged as a valuable approach to reduce both time and cost in drug discovery. Existing studies have primarily focused on predicting drug functions from one-dimensional and two-dimensional chemical structures and gene expression signatures. However, the potential of utilizing gene ontology, amino acid sequences, and drug-disease interactions in multi-label drug function classification has yet to be done. This study presents a methodology for identifying multiple drug functions based on therapeutic categories, leveraging gene ontology, amino acid sequences, and drug-disease interactions. Given the multi-label characteristic of drug function identification, a Binary Relevance (BR) approach is employed, along with four machine learning classifiers: K-Nearest Neighbour (KNN), Decision Tree (DT), Random Forest (RF), and Multi-Layer Perceptron Neural Network (MLPNN). The Multi-label Synthetic Minority Over-sampling Technique (MLSMOTE) is employed to tackle the prevalent class imbalance issue in multi-label datasets. Experimental results demonstrate promising performance for the present methodology, with the K-Nearest Neighbour classifier, achieving the highest accuracy of 99.23% in the drug-disease interactions dataset.

Read the paper · More papers on PaperTik