Two new molecular preprocessing schemes for machine learning and their evaluation using some DT algorithms

G. Vinotha, T.V. Sundar, D. Amuthalakshmi, M. Vivek · AIP conference proceedings · 2019

Machine learning is a popular branch of artificial intelligence to carryout exploratory data analysis and to solve various problems such as weather forecasting, drug discovery, encrypted image detection etc., In pharmaceutical industry, identification of a potential therapeutic drug for treating a disease involves the laborious initial process of screening hundreds of candidate molecules followed by the selection of a few for further processing. In such situations, Decision Tree (DT) algorithms could be helpful in identifying potential therapeutic drug molecules. The molecular data sets extracted in the form of numeric items or categorical entities or in mixed form can be fed to the DT algorithms as attributes to make the classification of the drug category or to predict the drug potency of molecules under test. In this regard, we have introduced two new schemes for data generation, one derivable from the geometry of the molecule and other derivable from the energy aspects of it. We have devised the schemes for some antiviral and antiparasitic drug molecules and used the generated data to evaluate the drug molecular screening efficiency of popular DT algorithms in the Waikato Environment for Knowledge Analysis (WEKA) platform.

Read the paper · More papers on PaperTik