Machine Learning for Imbalanced Data in Telecom Churn Classification
Olivia Intan Permata Dewi, Vina Nathania Santiko, Snow White Putri Safa, Alexander Agung Santoso Gunawan, Karli Eka Setiawan · 2024
In the highly competitive telecommunications industry, customer churn poses a significant threat to revenue and long-term sustainability. To proactively address this challenge, machine learning (ML) techniques offer promising solutions for predicting and mitigating customer churn. This study applies ML algorithms to churn prediction, focusing on the challenges of imbalanced data. Using a publicly available Telecom dataset from Kaggle, the data was meticulously prepared through a multi-stage process, including Random Under-Sampling (RUS) to address class imbalance. The study employed a range of ML algorithms, evaluating their performance using key metrics such as recall, precision, F1score, and accuracy. The findings reveal valuable insights into how RUS can enhance model performance by improving the prediction of the minority class. XGBoost outperformed Logistic Regression and Random Forest in accurately identifying positive outcomes, suggesting its effectiveness in recognizing true instances of customer churn. The confusion matrix supports this observation, showing that the recall value produced by XGBoost slightly exceeded those for Logistic Regression and Random Forest.