Phenotyping in distributed data networks: selecting the right codes for the right patients.
Anna Ostropolets, Patrick B. Ryan, George M. Hripcsak · PubMed · 2022
Observational data can be used to conduct drug surveillance and effectiveness studies, investigate treatment pathways, and predict patient outcomes. Such studies require developing executable algorithms to find patients of interest or phenotype algorithms. Creating reliable and comprehensive phenotype algorithms in data networks is especially hard as differences in patient representation and data heterogeneity must be considered. In this paper, we discuss a process for creating a comprehensive concept set and a recommender system we built to facilitate it. PHenotype Observed Entity Baseline Endorsements (PHOEBE) uses the data on code utilization across 22 electronic health record and claims datasets mapped to the Observational Health Data Sciences and Informatics (OHDSI) Common Data Model from the 6 countries to recommend semantically and lexically similar codes. Coupled with Cohort Diagnostics, it is now used in major network OHDSI studies. When used to create patient cohorts, PHOEBE identifies more patients and captures them earlier in the course of the disease.