ASIDL : A method for publishing missing data based on decomposition and slicing
Wu Tongtong, Zhang Hongxia · 2021
In data release, data mining and analysis technologies can fully tap the value of data and promote the development of related industries. However, maliciously mining data seriously threatens users' private information. Therefore, the privacy protection data release technology that realizes data security and effectiveness has developed rapidly, and corresponding research results have been obtained. However, the current research directions are mostly related to the general protection needs, and the data release area for special protection needs still needs attention. Aiming at the problem of reconstruction errors that are almost unavoidable in the reconstruction method of missing data, this paper proposes the ASIDL method based on decomposition and slicing. The TBBL (Tuple Bucket Build based on 1-diversity) algorithm clusters the QI attributes and randomly extracts the top 1 Tuples with different sensitive values form buckets, and more tuples formed by permutation and combination are used to protect their identity. At the same time, the TBSL (Tuple Bucket Split based on 1-diversity) algorithm is used to split the existing tuple buckets under the premise of satisfying 1-diversity, so as to achieve smaller buckets and ensure information validity. This paper compares with existing algorithms, and the results show that the method proposed in this paper has lower information loss and higher efficiency.