K-Anonymity without the Prior Value of the Threshold k: Revisited

Khan Farhan Rafat · 2019

k-Anonymity is a popular concept associated with a data set to designate it some level of anonymity. That is, a dataset is said to be k-anonymous if each of its identity-revealing attribute termed as Quasi-Identifier, appears for a minimum in at least k different tuples of the data set. Since its first publication in the year 2002, that concept has remained a focus of interest in the research arena, where the majority of studies have agreed on the prior delineation of the value of k. This research, however, adlibs one of the recent research that auto assigns to k an amount computed by traversing through each row in an anonymized data set. Our proposed solution suggests having a byte array of a dimension of 100, corresponding to the two-digit suppressed area code whose values can range from 0 to 99. The process requires the initialization of this array at the start with all zero/null values. With each row taken for pre-processing, that is, perform generalization and suppression, the value of the array element whose index corresponds to the value of the two-digit suppressed ZIP/Area Code is incremented by one. Whenever the resultant value exceeds 255, the same gets reinitialized with a value of two. The process terminates with the pre-processing of the dataset after which, a new kanonymized table is generated but only for the tuples whose ZIP/Area Code has a value > 2 in the corresponding index element of the proposed array. Doing so reduces the search iterations, which has a positive effect on process efficiency. Further, the proposed enhancement appreciates human intervention for add-in security during the publication process.

Read the paper · More papers on PaperTik