A Gabor Transform Method to solve the Automatic Tumor Classication Task

Michiel Moonen · 2009

A clinical or surgical pathologist is specialized in examining tissues and cells from the human body and body cavities, with the aim to grade them as benign or malicious. Performing the grading automatically by means of digital techniques is defined as the automatic tumor classification task. In an attempt to solve this, a digital database containing images of tissue is collected and labeled by an expert. Ficsor et al. [8] propose a method based on explicit information (explicit method) to determine parameters for automatic classification of three types of colitis (i.e. chronic inflammation of the colon). Explicit information is the information which can be explicitly described, e.g. color, size and shape. Although their research does not address the automatic tumor classification task, it is very relevant for comparison to our research since it automatically attempts to characterize and classify tissue. This thesis proposes and researches a method based on implicit information (implicit method): information that is implicitly available and not directly interpretable in terms of shapes or contours. Implicit methods are different from explicit methods because they rely on statistical measures of images, not necessarily confined to objects or contours. The problem statement of this thesis is defined as follows: Problem statement: How well does the Gabor transform (GT) method perform the automatic tumor classification task? The GT method comprises four stages. In the first stage the data is preprocessed, followed by the second stage where a log-Gabor transform is applied. In the third stage a statistical descriptor is applied (magnitude) and two types of features are extracted: the first type of feature is the maximum magnitude per scale, the second type of feature is the biased sum of histogram values per scale. These two features are dimensionally reduced by either Principal Component Analysis (PCA), Sammon Mapping (Sammon) or t-Stochastic Neighbor Embedding (t-SNE). In the fourth stage the features are presented for classification in four ways; besides dimensionally reduced by the three aforementioned techniques, it is presented as an unreduced feature vector. The data becomes classified by either Linear Discriminant Analysis (LDA), Support Vector Machine (SVM) or k-Nearest Neighbor (kNN). SVM is applied with one of two different scaling factors 1 and 13 (SVM-1 and SVM-13). Therefore, this thesis also adresses the following research questions: • Research Question 1: Which type of implicit feature performs best in the automatic tumor classification task? • Research Question 2: Which dimensionality reduction technique performs best in the automatic tumor classification task? • Research Question 3: Which classifier performs best in the automatic tumor classification task? By applying crossvalidation and confusion tables, the GT method was evaluated by its classification accuracy and false positive rate. Furthermore it is compared to the explicit method of Ficsor et al. We concluded that a combination of the maximum magnitude per scale, that is dimensionally reduced by t-SNE and classified by SVM-13 achieves the highest classification accuracy (87.5%), but still has a non-zero false positive rate (8.8%). For the automatic tumor classification task: • The best performing feature is the maximum magnitude per scale. • The best performing dimensionality reduction technique for the maximum magnitude per scale is t-SNE, and for the biased sum of histogram values per scale this is Sammon mapping, although dimensional reduction is not always necessary for this type of feature. • The best performing classifier for the maximum magnitude per scale is dependant on the data representation: if the dimensionally unreduced feature is applied, SVM-13 performs best. If PCA or t-SNE is applied, SVM-1 performs best. If Sammon mapping is applied, kNN performs best. For the biased sum of histogram values per scale, SVM-1 performs best one of three dimensionality reduction techniques is applied and LDA performs best when this type of feature is dimensionally unreduced. We conclude that the GT method performs the automatic tumor classification task very good compared to the explicit method by Ficsor et al.

Read the paper · More papers on PaperTik