Mean-entropy discretized features are effective for classifying high-dimensional bio-medical data

Jinyan Li, Huiqing Liu, Limsoon Wong · 2003

Abstract. This paper studies an empirical feature selection heuristics for classifying high-dimensional bio-medical data. A feature’s discriminating power can be measured by its entropy value. Based on this idea, we do not consider those features that are ignored by the entropy idea. Such a selection can usually reduce the dimensionality of the data by 90–95%. Then we rank the remaining features, and select features whose entropy is smaller than the average of all the remaining features ’ entropies. This round of selection can usually further reduce two thirds of the features. So, we can achieve a reduction from tens of thousands of features to only hundreds of important features. Furthermore, we also observe that learning algorithms, including our new tree-committee classifier, generally improve their accuracy after the feature selection. This heuristics appears to be more systematic than the prevailing use of specific numbers of top-ranked features for classification. 1

Read the paper · More papers on PaperTik