Bagging-Based Logistic Regression With Spark: A Medical Data Mining Method
Jian Pan, Yiang Hua, Xingtian Liu, Zhiqiang Chen, Zhaofeng Yan · 2016
Medical data in various organizational forms is voluminous and heterogeneous, it is significant to utilize efficient data mining techniques to explore the development rules of diverse diseases.However, many single-node data analysis tools lack enough memory and computing power, therefore, distributed and parallel computing is in great demand.In this paper, we propose a comprehensive medical data mining method consisting of data preprocessing and bagging-based logistic regression with Spark (BLR algorithm) which is improved for better compatibility with Spark, a fast parallel computing framework.Experimental results indicated that although the BLR algorithm took a little more duration than logistic regression (LR), it was 2.12% higher than LR in accuracy and outperformed LR with other common evaluation indexes.