Consensus Clustering And Fuzzy Classification For Breast Cancer Prognosis
Jonathan M. Garibaldi, Daniele Soria, Khairul A. Rasman · 2010
Extracting usable and useful knowledge from large and complex data sets is a difficult and challenging problem. In this paper, we show how two complementary tech-niques have been used to tackle this problem in the con-text of breast cancer. Diagnosis concerns the identifica-tion of cancer within a patient; in contrast, prognosis con-cerns the prediction of the ongoing course of the disease, including issues such as the choice of potential treat-ments such as chemotherapy or drug therapy, in combi-nation with estimation of chances (or length) of survival. Reliable prognosis depends on many factors, including the identification of the type of this heterogeneous dis-ease. We first use a consensus clustering methodol-ogy to identify core, well-characterised sub-groups (or classes) of the disease based on a large database of pro-tein biomarkers from over a thousand patients. We then use fuzzy rule induction and simplification algorithms to generate a simple, comprehensible set of rules for use in future model-based classification. The methods are de-scribed and their use is illustrated on real-world data.