Hybrid approach for tuberculosis data classification using optimal centroid selection based clustering
Manish Shukla, Sonali Agarwal · 2014
Application of classification technique in healthcare is challenging because of high dimensional medical data and of its dynamic nature. The research work here is focused on the study of various approaches for transformation large data into smaller datasets in effective manner so that accurate classification could be performed. Data clustering is a machine learning approach which divides dataset into smaller partitions and having higher intra partition similarity within it and dissimilarity among different partitions. Many clustering algorithm exists for varying nature of dataset and own their advantages as well as limitations as per nature of individual datasets thus there is sufficient scope to explore efficient and new algorithm for clustering based classification. This paper presents a new approach for centroid selection in k-mean algorithm for health datasets which gives better clustering results in comparison to traditional k-mean algorithm. The algorithm is evaluated against tuberculosis dataset and then results are applied to classifier for performance evaluation and results show improvement over previous algorithm.