Detecting Protected Health Information with an Incremental Learning Ensemble: A Case Study on New Zealand Clinical Text

Balkaran Singh, Quan Sun, Yun Sing Koh, Junjae Lee, Edmond Zhang · 2020

Clinical narratives host vast accumulations of patient data pivotal for research and development of health related products. In order for this data to be utilized, the underlying protected health information needs to be de-identified to ensure medical confidentiality. Given the voluminous size of clinical texts, manual de-identification of such large datasets is both expensive and impractical. Therefore, the concept of automated de-identification is a highly appealing prospect. Machine learning or model based sequential labeling algorithms, such as the named entity recognition algorithms and rule-based algorithms are among the most effective approaches to automated de-identification. A natural question to ask is how we can combine them to have the best of both worlds. In this paper, we present an analytical and easy to interpret framework to dynamically combine a sequential labeling model and a soft-rule-based model in an incremental learning setup. This framework is applied to a case study, which is part of a project prototyping automated de-identification system for New Zealand clinical free text data. Evaluations show that our approach can accommodate changes in the incoming data through dynamic updating. The simplicity of the framework also allowed us to gain insights on behaviour e.g. change of importance between the machine learning and rule models.

Read the paper · More papers on PaperTik