A study of performance enhancement in big data anonymization
Sung-Bong Jang · 2017
This paper presents the schemes to solve problems when k-anonymity and l-diversity are applied to Big-Data anonymization. The first problem is that information loss and distortion are unavoidable by anonymization job. To reduce the distortion, this paper presents an efficient method that is based on deep anonymization detection. In the method, data publishers analyze the anonymization work, and determine if it is deep or light. If it is thought as deep anonymization, high information distortion is allowed when being distributed to a third party after anonymization. Otherwise, information distortion is kept as low as possible when anonymizing Big-Data to provide the receivers with more meaningful data. The decision for deep anonymization is done by considering a domain data characteristic, data receiver's purpose, and data criticality. The second problem is that it takes much time and requires large buffer space to process the anonymization. To solve the problem, this paper present enhanced read/write schemes.