Data privacy and anonymization techniques in big data

Yueran Zhang · IET conference proceedings. · 2026

Data management, data privacy and anonymization is more important in the light of big data and people have to be cautious while dealing with such large datasets and need to be careful while following the legal requirements. This research work put forward a hierarchical model of anonymization procedures that combines state-of-the-art anonymization methods with machine learning strategies to ensure better privacy preservation along with data usefulness. The approach includes the steps of data gathering and preparation, data anonymization through differential privacy, k-anonymity, l-diversity and t-closeness models, and the utilization of artificial intelligence and federated learning methodologies. Experiments on real-world public datasets to analyze the applicability of the methodology conclude that it deliver rather satisfactory privacy protection and analytical utility. It consists of five types of privacy-preserving methods and shows that differential privacy and t-closeness provide the highest privacy and data utility while federated learning minimizes the risks of data leaks. While admitting that the proposed methodology is computational less efficient, especially in the process of t-closeness, it could be suggested that the overall perspective that has been provided to accomplish a balance between data privacy in big data settings is highly satisfactory. Finally, this work also suggests that the responsibilities for advancing and implementing privacy-preserving solutions include constant research and development with respect to advancing data privacy regulations that are compliance to.

Read the paper · More papers on PaperTik