A Hybrid Semi-supervised Learning Approach to Identifying Protected Health Information in Electronic Medical Records

Nguyen Dong Phuong, Vo Thi Ngoc Chau, Hồ Tú Bảo · 2016

De-identification of electronic medical records is one of the main tasks to make clinical data sharable for more researchers outside the associated institutions. Indeed, this de-identification task has been considered very much with positive research outcomes worldwide, especially those from the i2b2 (Informatics for Integrating Biology and the Bedside) shared tasks in 2006 and 2014. However, it has not yet been a solved problem and still needs more investigation realistically. In this paper, we propose a hybrid semi-supervised learning approach to identifying protected health information (PHI) in electronic medical records. The proposed approach combines a machine learning-based method with a conditional random fields (CRF) model and a rule-based method in a post-processing phase to handle 8 PHI types with disambiguity. The CRF-based classification phase and the rule-based post-processing phase are then conducted in a semi-supervised learning manner. As compared to the existing works, our work has the merits of PHI identification such as: (1). Effectiveness with comparable precision and recall values from the experiments on the 2006 i2b2 data set; (2). A more practical solution as enhancing the training data set over time for a more accurate classifier in a semi-supervised learning mechanism; (3). A portable approach for clinical text in non-English languages.

Read the paper · More papers on PaperTik