Steps in developing Watson for Oncology, a decision support system to assist physicians choosing first-line metastatic breast cancer (MBC) therapies: Improved performance with machine learning.
Julia Fu, Ayca Gucalp, Marjorie Glass Zauderer, Andrew S. Epstein, Mark G. Kris, Jeffrey Keesing, Aryeh Caroline, Mark Megerian, Thomas Eggebraaten, Robert DeLima, Marie Setnes, Patrick Wildt, Andrew David Seidman · Journal of Clinical Oncology · 2015
566 Background: IBM Watson for Oncology (WFO), trained by Memorial Sloan Kettering (MSK), is a cognitive computing system designed to apply machine learning, informed by MSK expertise, to enhance clinical decision making. We expanded the initial adjuvant therapy prototype to include MBC. Patients with MBC share many common attributes (e.g. age, performance status, receptor status, prior therapy) but receive widely varying therapies. Accuracy of WFO treatment recommendations for similar cases can be improved with machine learning features built on highly influential attributes. Methods: 101 manufactured MBC training cases were grouped by biologic characteristics and then analyzed for variation. A set of clinical attributes characterizing disease burden was posited to be influential in selecting the initial systemic therapy for MBC recommended by MSK breast medical oncologists (e.g. endocrine therapy vs chemotherapy for a patient with hormone receptor (HR)+MBC). WFO learned weights characterizing each attribute’s importance (e.g. site(s), extent, and size of metastatic lesions, degree of symptoms) through iterative experiments on training cases, and incorporated them into a “disease burden score”). Model performance as quantified by a combination of precision and recall (F1 Score) defined as (2 * true positives) / [(2 * true positives) + (false positives) + (false negatives)], was measured with and without the use of the disease burden score. Results: Activation of the disease burden score feature yielded a relative improvement in F1 Score of 11.5% across all cases (n = 101) from 73.6% to 82.1%; 28.8% among HR+ HER2+ cases (n = 32); 9.6% among HR+ HER2- cases (n = 29); 2.8% among HR- HER2+ cases (n = 17) and -1.4% among HR- HER2- cases (n = 23). Conclusions: Using training data and medical logic, WFO’s machine learning model can assign weights to attributes and features to select therapy for patients that better align with the nuanced decision-making of MSK breast medical oncologists than rules alone. WFO’s machine-learned disease burden score is a useful driver of treatment recommendations for HR+ and/or HER2+ MBC.