Defense Against Membership Inference Attacks on Imbalanced Dataset Using Generated Data
Qiaoling Lu, Feng Tian, Jie Ren, Mengdi Li · 2024
Machine learning models are vulnerable to membership inference attacks, where attackers aim to predict whether specific samples are present in the training dataset of the target model. Although the defense and attack frameworks for membership inference have been well developed, they mostly rely on balanced datasets. However, imbalanced datasets are prevalent, and algorithmic fairness imposes higher demands on our research. The impact of imbalanced datasets on model utility and membership inference attacks warrants further investigation, particularly concerning the effects on minority classes within these datasets. In this paper, we explore the effects of dataset imbalance on model utility and membership inference attacks. We propose mitigation methods for attacks on minority classes based on generated data tailored to imbalanced datasets. Experimental results demonstrate lower prediction accuracy and higher membership inference attack accuracy for minority samples. By addressing data imbalance through generated data, we ensure model utility while protecting the privacy of minority classes members against membership inference attacks.