Optimization of MDL Substructure Search Keys for the Prediction of Activity and Toxicity
Douglas R. Henry, Joseph L. Durant · ACS symposium series · 2005
This article describes the algorithmic generation of MDL substructure search keys and the optimization of the key definitions and weightings for predicting the biological activity of chemical structures. Substructure search keys are bitsets representing functional groups and atom pairs in molecules. They are mainly used for molecular similarity calculations, but they are also useful for clustering and classification analysis. We applied genetic algorithms to generate keysets that could better predict activity in a set of structures first used by Briem and Lessel (Briem, H. and Lessel, U., Perspect. Drug Discov. Design, 2000, 20, 231-244). Prediction performance improved from 65% to 74% correctly classified using a 324-keyset. We then applied a variety of weighting schemes in similarity calculations to discriminate between drug structures from the MDL Drug Data Report (MDDR) database and toxic structures from the MDL Toxicity database. The best results were obtained when the keys were weighted according to the inverse of database frequency (79% correct), followed by surprisal and unit weighting. Using coefficients from principal component and discriminant analyses did not yield better results.