A Multi-stage Ensemble Data Mining Model to Predict Ferritin Serum Levels
Mohammad Ali Abedini, Mohammadsadegh Mobin, Kamran Heidari, Afshan Roshani · 2015
Motivation: The Ferritin serum level is one of the key factors in diagnosing Iron Deficiency Anemia (IDA)-related diseases, which is one of the most common types of anemia. It is not common to measure the ferritin serum level in many cases, especially in the primitive stages of disease diagnostics; and in clinical laboratories it is not always feasible to assess ferritin serum levels. Objectives: In this research, we developed a multi-stage ensemble data mining model which predicts the ferritin serum level in a more efficient way. Summary of Proposed Model: The proposed model works as a Decision Support System (DSS) which considers the Complete Blood Count (CBC) test results as inputs in order to make a prediction for ferritin serum levels. The developed model uses demographical information of the patients in addition to CBC test results consisting of three stages: 1. Select important features using correlation-based feature selections; 2. Train the decision tree as a base classifier by applying four different ensemble regressions approaches including: Bagging, Additive regression, Rotation forest and Random subspace; 3.Evaluate and compare mentioned approaches based on correlation coefficient and root mean squared error criteria. Summary of the Results: The results show that the bagging approach outperforms the others in terms of both criteria. By conducting this case study, the proposed model has proven to be an efficient DSS in IDA diagnosis. Dataset Description Dataset was obtained in TALEGHNI Hospital, Tehran, Iran. About 300 people were selected from the hospital patient list and after initial assessments by Department of Clinical Diagnostics; a dataset size of 164 was selected. Dataset Partitioning Dataset: 164 Train Set: 114 Test Set: 50 Model Input & Output Inputs (Features) ● CBC Test Result: 1. Red Blood Cells(RBC) 2. Hemoglobin (HG) 3. Hematocrit (HCT) 4. Mean Corpuscular Volume (MCV) 5. Mean Corpuscular Hemoglobin (MCH) 6. Mean Corpuscular Hemoglobin Concentration (MCHC) ● Demographical Characteristics: 7. Age 8. Sex Output: Ferritin Serum Level Note: All features from CBC Test and Ferritin are numeric. PROPOSED DATA MINING MODEL Data Mining Software Weka is a collection of Machine Learning Algorithms for data mining tasks. It contains tools for data pre-processing, classification, regression, clustering, association rules, and visualization. It is also well-suited for developing new machine learning schemes. The Proposed Model Framework 1. Feature Selection ● Correlation-based feature selection ● Search method: Best first 2. Ensemble Learning Findings: Input Features for the Regression Model ● Base Regression: REPTree ● Ensemble methods: 1. Bagging 2. Additive regression 3. Rotation forest 4. Random subspace Findings: Four Different Trained Ensemble Regression models