A framework for a privacy-aware feature selection evaluation measure

Yasser Jafer, Stan Matwin, Marina V. Sokolova · 2015

Feature selection is based on the notion that redundant and/or irrelevant variables bring no additional information about the data classes and can be considered noise for the predictor. As a result, the total feature set of a dataset could be minimized to only few features containing maximum discrimination information about the class. Classification accuracy is used as the evaluation measure in guiding the feature selection process. At the same time, such measure does not take into account the privacy of the resulting dataset. In this work, we incorporate privacy considerations into the very evaluation measure that is used to evaluate and select feature subsets. We consider privacy “during” the feature selection process and as such introduce a two-dimensional measure in automatic feature selection that takes into account both objectives of privacy and efficacy (e.g. accuracy) simultaneously and provides the data user with the flexibility of trading-off one for another.

Read the paper · More papers on PaperTik