A Comprehensive Review of Machine Learning Privacy
Haoru Chen · 2024
The integration of machine learning (ML) into various domains has raised significant concerns regarding the privacy of individuals. As datasets grow larger and more complex, the potential for sensitive information leakage during the learning process becomes a critical issue. This paper provides a comprehensive review of the current state of machine learning privacy, examining the theoretical foundations, practical challenges, and emerging solutions. It delineates the context of ML privacy issues, focusing on the data-centric approach of ML, its wide-ranging applications, and the resultant privacy exposures. The paper thoroughly describes common privacy attack methodologies, which include membership inference, model inversion, and poisoning attacks. The review then catalogues current privacy-preserving techniques, encompassing encryption strategies, differential privacy, and homomorphic encryption. It evaluates privacy safeguarding approaches designed for distinct ML environments, spanning centralized, distributed, and federated learning architectures. Ultimately, the paper forecasts the trajectory of future ML privacy research, highlighting the imperative for innovation in technology, collaborative interdisciplinary work, and the refinement of legal and regulatory frameworks to enhance privacy in ML contexts.