Estimating the Statistical Significance of Classifiers used in the Prediction of Tuberculosis
T Asha · IOSR Journal of Computer Engineering · 2012
Tuberculosis (TB) is a disease caused by bacteria called Mycobacterium Tuberculosis.It usually spreads through the air and attacks low immune bodies.Human Immuno deficiency Virus (HIV) patients are more likely to be attacked with TB.It is an important health problem in India as well.Diagnosis of pulmonary tuberculosis has always been a problem.Classification in medicine is an important task in the prediction of any disease.It even helps doctors in their diagnosis decisions.However the decision of best classification cannot just depend on accuracies or error rates.There is a need for critical statistical analysis of these classifiers based on some statistical tests.In this paper, a study on classification of Tuberculosis with statistical significance is realized at two stages.First stage is the comparison of accuracies by classifying TB data into two categories Pulmonary Tuberculosis(PTB) and retroviral PTB(RPTB) ie TB along with AIDS using basic learning classifiers such as C4.5 Decision Tree, Support Vector Machines (SVM), K-nearest neighbor, Bagging and Naïve Bayesian algorithms.Second stage is evaluating the performance of these classifiers using paired ttest to select the optimum model.Results for our datasets show that SVM and C4.5 Decision Tree are not statistically significant, whereas SVM with Naïve Bayes and K-nearest neighbor are statistically significant.