Embedded Cluster Modelling-A novel method for analysing embedded data sets

Andrew Paul Worth, M Cronin · Quantitative Structure-Activity Relationships · 1999

Cluster Significance Analysis (CSA) is a method for analysing embedded data sets, i.e. data sets in which the objects (chemicals) are divided into two classes (active/inactive or toxic/non-toxic) and in which one class of objects (typically, the active or toxic chemicals) is found to cluster along one or more variables (e.g. physicochemical descriptors), forming an ‘embedded cluster’ surrounded by the ‘diffuse cluster’ of objects in the other class (typically, the inactive or non-toxic chemicals). The aim of CSA is to identify variables along which clustering is statistically significant. Having identified significant variables, the investigator may wish to derive a model for classifying active and inactive chemicals on the basis of these variables. In this paper, a method called ‘embedded cluster modelling’ (ECM) is proposed for the derivation of such classification models. If ECM is applied to a single variable, the resulting model consists of two cut-off values (an upper and a lower limit) between which the active (toxic) chemicals are predicted to lie. If ECM is applied to two or more variables, the resulting model is best described as an ‘elliptic model’ of cluster membership, since the active (or toxic) chemicals are predicted to lie inside the boundary of a two-dimensional or three-dimensional ellipse, which is regarded as the boundary of the embedded cluster. The combined use of CSA and ECM for the analysis of embedded data sets is illustrated by their application to a data set of methacycline derivatives. The algorithms for CSA and ECM have been coded in the form of Minitab macros, which the authors are making freely available.

Read the paper · More papers on PaperTik