Multi-objective clustering of gene expression data with evolutionary algorithms: a query gene approach
Michael Calonder · Repository for Publications and Research Data (ETH Zurich) · 2006
Biologists are interested in the discovery of co-regulated genes since such genes are likely to share a common biological function.This work describes an evolutionary algorithm to find groups of genes in expression data that exhibit expression profiles similar to that of a given query gene.We are clustering over multiple data sets, such that the task is to maximize the overall co-expression.To this end, we treat the homogeneity of a certain gene group in every data file as an independent objective and employ a multi-objective evolutionary algorithm.In this context we also discuss three similarity measures, namely Euclidian distance, Pearson correlation and the ranked mean squared residue measure.We found that there is a clear trade-off between the abovementioned objectives.The query gene clustering algorithm is validated by showing that it is able to recover an already known cluster with 32 genes.In a second step we introduce the cluster size as another objective that we require to be maximized.We found that this additional objective leads to an improvement in the quality of the clustering result and document this finding by 40 cases with different parameter settings.