Analysis and Prediction of Drugs using Machine Learning Techniques
Abhay Dadhwal, Meenu Gupta · 2021 3rd International Conference on Advances in Computing, Communication Control and Networking (ICAC3N) · 2021
In the era of increasing demand of data analytics in drug discovery domain, processing the complex high dimensional drug datasets is a challenging research area. High dimensionality refers to dataset having more number of columns than rows in which there may be irrelevant features that are futile. For reducing high dimensionality issue, feature selection techniques are widely adopted. Feature selection aims to identify relevant subset of features from original feature set before the training the machine learning model. From variety of feature selection techniques available one can select according to his requirement that can include faster processing, low resource consumption, high accuracy etc. In this paper, a high dimensional drug dataset is used to train the machine learning model. The rationale is to analyze and compare the performance of various feature selection techniques namely Pearson Correlation Coefficient, Chisquare test, Lasso, Treebased, Anova-f to produce best predictive performance. When a feature selection method is applied, the results are tested using K fold cross validation with state-of-the art machine learning algorithms. These numerical results of methods are presented and compared to acquire insight about these methods leading to better generalization and accuracy and also in making better future decisions.