Quantification of Molecular Similarity and Its Application to Combinatorial Chemistry
Richard A. Lewis, Andrew C. Good, Stephen D. Pickett · 1997
The concept of molecular similarity is an essential part of modern computeraided drug design methods, and has been successfully applied many times in the optimisation of lead series. Molecular diversity calculations are becoming an increasingly important tool in the field of combinatorial chemistry. For many research groups, this new chemical technology is and will continue to produce a huge increase in the number of compounds available for screening. In many cases, the number of molecules that could be constructed greatly exceeds potential synthesis and screening capacity; it will become vital to design libraries based on the properties of compounds already in existence, if the diversity of each new molecular collection is to be maximized. The central requirements of such an approach are the construction of a product-based descriptor which can be applied to an existing compound data set, and a robust method for optimising the choice of reagents to maximise diversity. This presents a major problem for many of the paradigms currently used in such calculations, as they are reagent-based and thus not suitable as descriptors of whole molecules. In this paper, we will discuss briefly some of the approaches and metrics used to describe molecular similarity, and techniques for combinatorial library design. We will then describe a novel product-based method for determining molecular diversity, comparing it critically to the ChemDiverse software. The new method includes direct reagent selection, sophisticated diversity measurement functions and improved pharmacophore keys. The potential of this tool to aid in diversity profiling is illustrated in three studies. We will finish with a discussion of the experimental validation of similarity measures, with the conclusion that this goal is not readily attainable given current information.