Proposal of l-diversity algorithm considering distance between sensitive attribute values
Keiichiro Oishi, Yasuyuki Tahara, Yuichi Sei, Akihiko Ohsuga · 2017
Consideration of privacy is crucial when sharing a database that contains personal information with other organizations. Many organizations have utilized personal information while realizing the importance of personal privacy protection by anonymizing personal information according to existing indicators, such as k-anonymity. A database with personal information is defined as satisfying l-diversity when a specific record group that has the same combination of quasi-identifiers (QIDs) holds at least / kinds of sensitive attribute value. By satisfying l-diversity, the identification probability of the individual's sensitive attribute value becomes less than 1/l, and it can be said that privacy is protected. The l-diversity has been widely studied in the area of privacy-preserving data mining. However, if a database containing certain personal information holds similar sensitive attribute values, there is a possibility that de facto diversity is not satisfied, even if anonymization is performed to satisfy l-diversity. In this research, we propose (l, d)-semantic diversity that is able to consider more actual diversity to solve the problem of not being able to satisfy de facto diversity with the existing indicator. The (l, d)-semantic diversity considers the similarity of sensitive attribute values by adding distances, d, defined using categorization. We also propose an anonymization algorithm and analysis algorithm suitable for the proposal indicator, and we conduct evaluation experiments.