Efiective Linear Discriminant Analysis for High Dimensional, Low Sample Size Data
Zhihua Qiao, Lan Feng Zhou, Jianhua Z. Huang · 2007
In the so-called high dimensional, low sample size (HDLSS) settings, LDA possesses the \data property, that is, it maps all points from the same class in the training data to a common point, and so when viewed along the LDA projection direc- tions, the data are piled up. Data piling indicates overfltting and usually results in poor out-of-sample classiflcation. In this paper, a novel approach to overcome the data piling problem is introduced. It incorporates vari- able selection into LDA. The underlying assumption is that, among the large number of variables there are many irrelevant or redundant variables for the purpose of classiflcation. By using only important or signiflcant variables we essentially deal with a lower dimensional problem. Experiments on both synthetic and real data sets show that the proposed method is efiective in overcoming the data piling and overfltting problem of LDA while improving the out-of-sample classiflcation performance. An important query in application of Fisher's LDA is whether all the variables on which measurements are ob- tained contain useful information or only some of them may su-ce for the purpose of classiflcation. Since the variables are likely to be correlated, it is possible that a subset of these variables can be chosen such that the others may not contain substantial additional informa- tion and may be deemed redundant in the presence of this subset of variables. A case for variable selection in Fisher's LDA can be made further by pointing out that by increasing the number of variables we do not necessar- ily ensure an increase in the discriminatory power. This is a form of overfltting. One explanation is that when the number of variables is large, the within-class covariance matrix is hard to be reliably estimated. In additional to avoiding overfltting, interpretation can be facilitated if we incorporate variable selection in LDA. We flnd that variable selection may provide a promising approach to deal with a very challenging case of data mining: the high dimensional, low sample size (HDLSS, Marron et al., 2007) settings. The HDLSS means that the dimension of the data vectors is larger (often much larger) than the sample size (the number of data vec- tors available). HDLSS data occur in many applied areas such as gene expression microarray analysis, chemomet- rics, medical image analysis, text classiflcation, and face recognition. As pointed out by Marron et al. (2007), clas- sical multivariate statistical methods often fail to give a meaningful analysis in HDLSS contexts.