A Breast Cancer Risk Classification Model Based on the Features Selected by Novel F-Score Index for the Imbalanced Multi-Feature Dataset

Xiaoli Lin, Wei Huangfu, Fei Wang, Liyuan Liu, Keping Long · 2016

Breast cancer has become the highest incidence of malignant tumors to global women. The breast cancer risk classification model can help reduce the incidence rate of breast cancer. For the large population of Chinese women, it is important to build an apt classification model and only the respondents in the high risk group will accept further diagnosis to filter out the breast cancer patients to lower the total medical cost. The classification model must have a low false-negative rate and also be low-cost. A novel one-class F-score index is introduced in this paper. We construct the naive Bayesian classifier based on the selected features inspired by the one-class F-score values. The experiment results show that, with the presented method, the false-negative rate is decreased to 0.039 with only 15 features. Compared with related methods, our method leads to the lowest false-negative rate and the lowest number of selected features. This work hopes to be a first step toward further practical breast cancer risk classification model for Chinese women.

Read the paper · More papers on PaperTik