Design of Small Libraries for Lead Exploration

Per M. Andersson, Anna Linusson, Svante Wold, Michael Sjöstróm, Torbjörn Lundstedt, Bo Nordén · Kluwer Academic Publishers eBooks · 2005

A combinatorial chemical library is a (usually large) set of compounds made to contain all possible structures of a certain type. The library is often made in order to find a lead compound for a specific drug action or for the optimisation of a lead. Because of the large number of synthesised compounds in the library, their biological activity is usually measured by rapid and simple tests, i.e. “high throughput screening” (HTS), giving crude answers, for instance “active” or “not”. Combinatorial chemistry (CombC) comprises a chain of parts linked by the objective of finding lead compounds for further development. Sometimes the objective is to optimise an existing lead compound, but this is not much discussed in this chapter. An analysis of this CombC chain indicates that the biological testing is the weakest part of the chain. This is due to the difficulty in performing an in-depth biological testing of any set of compounds exceeding a couple of hundred members. Hence there is a strong motivation to decrease the size of libraries to a size that allows in-depth biological testing.We discuss how the size of a library can be drastically reduced without loss of information or decreases in the chances of finding a lead compound. The approach is based on the use of statistical molecular design (SMD) for the selection of library compounds to synthesise and test, followed by the use of quantitative structure activity relationships (QSARs) for the evaluation of the resulting test data.The use of SMD and QSAR is, in turn, critically dependent on an appropriate translation of the molecular structure to numerical descriptors, the recognition of inhomogeneities (clusters) in both the structural and biological data spaces, and the ability to analyse and interpret relationships between multidimensional data sets.We present a strategy for constructing a library with optimal information while still taking synthetic feasibility into account. The objective is to provide optimal chemical diversity with a moderate number of compounds, plus adequate depth and width of the biological testing. The strategy is based on a multivariate characterisation of the synthesis starting materials (building blocks), Principal Component Analysis (PCA), multivariate design, and Multivariate Quantitative Structure-Activity Relationships (M-QSAR). The strategy applies both to solid phase synthesis as well as libraries in solution.

Read the paper · More papers on PaperTik