Privacy-Preserving Data Mining Techniques in Big Data Environments

M Anoop, G Michael, Jeni Gracia RC · 2025

The rapid increase in the data in every room has fueled the implementation of big data analytics. There is, however, a significant concern of privacy brought about by the use of such big, sensitive data. The challenge is solved with the help of Privacy-Preserving Data Mining (PPDM) that allows obtaining important insights without having to sacrifice individual confidentiality. The proposed paper above is a literature survey and critical analysis of PPDM techniques in big data environment. It takes the form of a categorization and then comparison between the privacy paradigms: anonymization, differential privacy, secure multiparty computation, and homomorphic encryption, and it investigates the strengths and trade-offs between privacy and utility they offer. The characteristics of big data namely volume, velocity, variety, and veracity are given particular importance along with the degree to which the current PPDM solutions have to deal with them. The proposed framework to evaluate the comparative situation is with experimental verification on real-world datasets conducted to evaluate the strength of privacy, accuracy, and scalability. Findings show that although such techniques as differentially private provide high levels of protection, they may affect data utility and time when processing data. The paper ends with the (inescapable) prediction that adaptive, scaleable and compliance aware PPDM solutions are needed in the future. The study highlights the necessity to further develop dynamic privacy solutions, the integration of federated learning, and compliance with laws to introduce safe, ethical data mining practices in the scenario of big data.

Read the paper · More papers on PaperTik