Model evaluation of datasets using critical dimension model invariants

Divya Suryakumar, Andrew H. Sung, Subhasish Mazumdar, Qingzhong Liu · 2012

Critical dimension is the minimum number of features that is required to ensure the performance of a learning machine to be “high”. This critical dimension is usually unique to the learning machine and the ranking algorithm combination. Medical- and bio-informatics datasets are different from most other datasets in that there is an imbalance in most of these datasets and a high prediction accuracy often depends upon not just the overall accuracy but also the true positive and the false negative rates. To find a medically and bio-informatically accurate critical dimension and for better analysis of such datasets we develop two evaluation models, one using all features and the other using critical number of features. The performance measurements such as accuracy, specificity, sensitivity, area under the curve, F-score and kappa values are compared. This paper shows that at the critical dimension the evaluation model shows good results for all performance measurements measured on most datasets studied. The difference in performance measurements obtained using only critical number and using all features is significantly less, i.e., there is not much difference in sensitivity, specificity and other measurements calculated.

Read the paper · More papers on PaperTik