Mitigation Strategies for Data Leakage in Machine Learning System Design
Manan Gupta, Jabez J. Christopher, R Kavya · 2025
Data leakage occurs when a Machine Learning (ML) model attains knowledge outside of training data. It undermines the validity of a model because it causes the model to learn from information it wouldn't have in real-world applications, leading to overestimated performance metrics and poor generalization to unseen data. Among different types of data leakages, this work focuses on train-test contamination which refers to the overlap between training and testing datasets, and target leakage which refers to the set of features that replicate the behaviour of target variable. Two strategies, overlap dropping and feature dropping, are proposed to mitigate train-test contamination and target leakage, respectively. Overlap dropping discards contaminated testing samples, while feature dropping eliminates replicated features. Extensive experiments were conducted on publicly available benchmark datasets to evaluate these strategies. The findings revealed that no dataset inherently caused leakage in models. Additionally, overlap dropping and feature dropping effectively demonstrated the impact of similarity and correlation on model performance. Therefore, these actionable data leakage mitigation strategies can enhance the reliability and integrity of data science practices.