Privacy preservation for medical dataset using Hadoop

Balaji K. Bodkhe, Sanjay P. Sood · 2017

With the widespread of computing services all over the world and increase in technological dependency, medical data serves as a huge contributor to Big Data. As the data generation in the healthcare field is increasing, privacy preservation of this data presents itself as a challenge, hence, the need to protect the identity of a person is becoming more critical. In this project, we are implementing privacy preservation of Big Data using the Ha do op Analytic To ol: HIVE. This tool helps in parallel pro cessing of the data which makes the system efficient. HIVE facilitates querying and managing large datasets residing in dis tributed storage. We implement privacy preservation on a healthcare dataset hence preserving the identity of a person and their related diseases (sensitive attribute). We are using techniques and algorithms such as generalization, bucketization and suppression and slicing. These algorithms ensure Privacy preservation and maintain the data utility. Along with this we are also analyzing healthcare data in order to draw patterns with respect to certain attributes such as area and ill ness. These patterns will then be shown to authorize users.

Read the paper · More papers on PaperTik