Unraveling Model Inversion Attacks: A Survey of Machine Learning Vulnerabilities
Tanzim Mostafa, Mohamed I. Ibrahem, Mostafa M. Fouda · 2024
Model Inversion (MI) attacks are a class of security threats targeting machine learning (ML) systems. These attacks aim to uncover sensitive information that was used to train the ML model, by either reconstructing the original training data or inferring private attributes based on the model’s outputs. This research offers a comprehensive survey of MI attacks with a focus on their operation procedure, real-world implications, and defense strategies against such attacks. Additionally, it examines all MI attack survey papers to date and distinguishes itself through the most extensive and comprehensible categorization of existing MI techniques. Lastly, it highlights the overarching need for establishing strong defensive mechanisms to ensure protection against such advancing threats in cybersecurity.