Using Similarity Metrics on Real World Data to Recommend the Next Treatment

Kyle Haas, Malika Mahoui, Simone Gupta, Stuart Duncan Morton · 2018

Studies using similarity metrics have been used to help quantify the relationship between patients; however, these studies do not leverage either the patients' prior treatments or the ordering of these treatments. Our proposal seeks to recommend the next treatment for a given patient by comparing the overall survival of similar patients who share a common treatment stem. Data was aggregated from the FlatIron® Advanced non-small-cell lung (NSCLC) proprietary dataset [1] comprised of 1312 patients from 2008-2016. Our methodology pipeline was comprised of three main components (non-treatment-based similarity (NTS), treatment-based similarity (TS), recommendation). For NTS, we divided all non-treatment features into 2 main categories (i.e. genetic category and clinical/demographic category), computed a patient similarity using Gower Similarity Metric [2] for each category, and lastly, followed a similar approach as Gottlieb et al [3] and created a single similarity measure using the geometric mean. A similarity threshold is used to select for a reference patient p a set of similar patients Simp with a similarity value (determined from the genetic and clinical/demographic features) above the threshold value. TS is used during the next step to filter from Simp the patients that do not share the same treatment class-level stem (prior treatments) as the reference patient. The objective is to consider only patients who share similar previous class treatments in order to determine the next treatment class for the reference patient. From this final subset of patients, we determine which patient has survived the longest number of days following the treatment stem and we recommend this patient's next treatment class to the reference patient. To evaluate this approach, we repeated this methodology across 10 random subsets of patients where each subset was 10% of the entire data set. Each patient in each subset was viewed as a reference patient and compared against the patients outside the random subset. We varied the length of the initial treatment stem and varied the NTS similarity threshold value. We found that for stems lower than two treatments and with similarity thresholds above 0.6, only approximately 30% of patients took the same treatment as the longest surviving patient in the subpopulation. Further work is needed to refine the proposed approach, including assigning different weights to the features used in the similarity computation, and considering other outcome variables to recommend next treatment (e.g. quality of life using ECOG performance score).

Read the paper · More papers on PaperTik