Effects of data anonymization on the data mining results

Ines Buratovic, Mario Miličević, Krunoslav Žubrinić · International Convention on Information and Communication Technology, Electronics and Microelectronics · 2012

This article examines the possibility of publication of students' data, such as secondary school success, state graduation exam scores and success during their first year of university study for analyses. In order to discover data patterns and relationships using data mining techniques, the data must be released in the form of original tuples, instead of pre-aggregated statistics. These records contain sensitive and even confidential personal information, which implies significant privacy concerns regarding the disclosure of such data. Removing explicit identifiers prior to data release cannot guarantee anonymity, since the datasets still contain information that can be used for linking the released records with publicly available collections that include students' identities. One of the privacy preserving techniques proposed in the literature is the k-anonymization. The process of anonymizing a data set usually involves generalizing data records and, consequently, it incurs loss of relevant information. In the primary research undertaken in the University of Dubrovnik's students' database the effect of anonymization has been measured by comparing the results of mining the original data set with the results of mining the altered data set to determine if it is possible to use anonymized data for research purposes.

Read the paper · More papers on PaperTik