Privacy-Preserving Inductive Learning with Decision Trees

Stacey Truex, Ling Liu, Mehmet Emre Gürsoy, Lei Yu · 2017

With the continued explosion of digitized data, data mining and data collection have become more prevalent. With this growth, we have also seen increased concern over data privacy and intellectual property. Within this environment, an important question has emerged: Can machine learning and data mining techniques be leveraged without compromising privacy? This paper revisits the concepts and techniques of privacy-preserving decision tree learning, a fundamental model of inductive learning. We first examine different privacy risks during decision tree based inductive learning processes, including the sensitivity of private input data and potential privacy risks induced by inference over both the learning output and the intermediate results of inductive learning iterations. We then review and compare the privacy notions and properties of three orthogonal and yet complimentary technical frameworks: randomization-based data obfuscation, differential privacy, and secure multiparty computation. We analyze their effectiveness and review representative approaches in each of these three frameworks. Finally, we highlight some of the open challenges to privacy-preserving solutions for decision tree learning.

Read the paper · More papers on PaperTik