Privacy-preserving Data Splitting Based on Machine Learning

Huafeng Ruan · 2023

Privacy-preserving data splitting is an effective means of storing parts of data across different servers to safeguard user privacy. Current privacy protection schemes that employ data splitting for privacy protection assume that data owners are already aware of the constraints between attributes, which limits their applicability. To address this challenge, we propose a scheme that utilizes machine learning to segment anonymous data and securely store it. Initially, we employ machine learning algorithms to recursively compute the relationships between attributes. Subsequently, we introduce a greedy algorithm that leverages the relationships derived in the previous step to calculate a data fragmentation approach that ensures privacy safety. Experimental results indicate that our scheme can successfully derive privacy-safe fragmentation methods using various machine learning algorithms. Among the four machine learning algorithms tested, Lasso regression demonstrated the fastest computation speed.

Read the paper · More papers on PaperTik