Evaluation of Learning from Screened Positive Examples

Marina V. Sokolova, Khaled El Emam · 2008

In this work, we propose metrics to assess the classification performance of an algorithm that learns from a very small number of examples of one class. The metrics, True Detection Probability and False Referral Probability , have been used in medicine to evaluate classification performance from examples screened to be positive. Necessity of such metrics comes from a real life application: we are developing a tool to detect Personal Health Information (PHI) in files obtained from peer-to-peer file sharing networks. On the data available to us, the files with PHI represent a small portion of all files (1%). However, the detection and prevention of PHI leaks is important: such inadvertent disclosure of PHI increases opportunities for privacy breaches. 1. Appraisal of Machine Learning

Read the paper · More papers on PaperTik