An evaluation of feature sets and sampling techniques for de-identification of medical records
James Gardner, Li Ping Xiong, Fusheng Wang, Andrew Post, Joel Haskin Saltz, Tyrone W. A. Grandison · 2010
De-identification of text medical records is of critical importance in any health informatics system in order to facilitate research and sharing of medical records. While statistical learning based techniques have shown promising results for de-identification purposes, few such systems are publicly available. It remains a challenge for practitioners to build an accurate and efficient system as it involves a significant amount of feature engineering, i.e. creation and examination of new features used in the system. A comprehensive evaluation is needed to thoroughly understand the effects of different feature sets and potential impacts of sampling and their trade-offs between the often conflicting goals of precision (or positive predictive value), recall (or sensitivity), and efficiency.