Privacy Preserving Data Mining: A Novel Approach to Secure Sensitive Data Based on Association Rules
Syed Shujauddin Sameer · 2014
Abstract— The availability of data on the internet is increasing on a larger basis daily .Privacy Preservation data mining has emerged to address one of the side effects of data mining Technology. The threat to individual privacy through data mining is able to infer sensitive information from Non-sensitive information or unclassified data. There is a n urgent need to be able to infer some mechanism to avoid the projection of all the sensitive information .An approach in data mining techniques is very much essential. Alteration of data, filtering of the data, blocking of the data are Some of the approaches. Given specific rules to be hidden, the techniques involve is to hide only the given sensitive data. In this work we assume that only sensitive datais given and we analyze the approaches to secure sensitive data in the database. A key problem, and still not sufficiently investigated, is the need to balance the confidentiality of the disclosed data with the legitimate needs of the data users. Every disclosure limitation method affects, in some way, and modifies true data values and relationships.In this paper, we investigate confidentiality issues of a broad category of rules, which are called association rules. Sometimes, sensitive rules should not be disclosed to the public since, among other things, they may be used for inferencing sensitive data, or they may provide business competitors with an advantage.Many government agencies, businesses and non-organizations in order to support their short and long term planning activities, they are searching for a way to collect, analyze and report data about individuals, households or businesses. Information systems, therefore, contain confidential information such as social security numbers, income, credit ratings, type of disease, customer purchases, etc (10). Ideally, these effects can be quantified so that their anticipated impact on the completeness and validity of the data can guide the selection and use of the disclosure limitation method.