An Interval Tree Based Feature Reduction Method For Cancer Classification Using High-Throughput DNA Copy Number Data.

Siling Wang, Yuhang Wang, Luc Girard, Young Kim, Jonathan R. Pollack, John Dorrance Minna · 2007

data is an important bioinformatics problem. Effective machine learning models for this task can be useful not only for cancer diagnosis, but also for discovering novel tumor suppressor genes and oncogenes. The recent array-based assays that detect DNA copy numbers contain very large numbers of probes and thus generate data of extremely high dimensionality. Therefore, the use of appropriate feature reduction methods is called for. In this paper, we proposed an efficient interval tree based feature reduction method for cancer classification using DNA copy number data. Instead of using probes as features, our approach extracts intervals as features from the original probe data. Experiment re-sults on two real data sets showed that our approach led to statistically significantly better classification accuracies as compared to the based line approach where the DNA copy number data at probe loci were used as features directly.

Read the paper · More papers on PaperTik