Data Privacy in Machine Learning: A Pipeline for Privacy Risk Assessment
Epifelward Niño O. Amora, Michelle P. Ombid · 2025
Data privacy remains a critical concern in the development of machine learning (ML) systems, particularly when such models are trained on demographic datasets containing quasi-identifiable information. This study introduces a novel Privacy Risk Score, a feature engineered to quantify the re-identifiability of individual records based on the frequency of quasi-identifier combinations. Using the UCI Adult dataset, we implement a five-stage ML pipeline encompassing data collection, cleaning, feature engineering, model selection, and evaluation. A variety of classifiers including Logistic Regression, Support Vector Machine, Random Forest, and XGBoost are assessed for performance, with XGBoost achieving the highest test accuracy at 92.8%. Comparative experiments with and without the Privacy Risk Score reveal its significant impact on model performance, as visualized through ROC curves and confusion matrices. The findings underscore the importance of embedding privacy-aware features into ML pipelines, offering a technically rigorous and ethically responsible approach for handling sensitive personal data in predictive systems.