PD58-09 EXTRACTING STRUCTURED INFORMATION FROM PATHOLOGY REPORTS USING NATURAL LANGUAGE PROCESSING AND MACHINE LEARNING

Anobel Y. Odisho, Briton Park, Nicholas A. Altieri, William J. Murdoch, Peter Carroll, Matthew Coopberberg, Bin Yu · The Journal of Urology · 2019

You have accessJournal of UrologyGeneral & Epidemiological Trends & Socioeconomics: Quality Improvement & Patient Safety IV (PD58)1 Apr 2019PD58-09 EXTRACTING STRUCTURED INFORMATION FROM PATHOLOGY REPORTS USING NATURAL LANGUAGE PROCESSING AND MACHINE LEARNING Anobel Odisho*, Briton Park, Nicholas Altieri, William Murdoch, Peter Carroll, Matthew Coopberberg, and Bin Yu Anobel Odisho*Anobel Odisho* , Briton ParkBriton Park , Nicholas AltieriNicholas Altieri , William MurdochWilliam Murdoch , Peter CarrollPeter Carroll , Matthew CoopberbergMatthew Coopberberg , and Bin YuBin Yu View All Author Informationhttps://doi.org/10.1097/01.JU.0000557177.97226.63AboutPDF ToolsAdd to favoritesDownload CitationsTrack CitationsPermissionsReprints ShareFacebookLinked InTwitterEmail Abstract INTRODUCTION AND OBJECTIVES: Detailed pathologic information for the estimated 1.74 million Americans diagnosed with cancer every year is locked away as unstructured free text, unavailable for use without manual abstraction. Our objective was to optimize NLP algorithms to extract detailed pathologic details from cancer pathology reports. Building on current pipelines, we developed feature specific optimizations for information extraction from prostate cancer pathology reports and evaluate if high quality information extraction was possible with a minimal training data. METHODS: We used a corpus of 3,232 free text pathology reports from radical prostatectomy specimens at UCSF, each with detailed manual annotations for 20 data elements, such as gleason scores, margin status, extracapsular extension, seminal vesicle invasion, tumor volume, and numbers of lymph nodes (positive and dissected). The full corpus was divided up so that the training, validation, development test, and true test sets contained 65%, 15%, 10%, and 10% reports each, respectively. We then created an NLP pipeline using NLTK and investigated the performance of multiple machine learning methods using scikit-learn and pytorch. We applied random forests, support vector machines, boosting, logistic regression and convolutional networks to the full training dataset as well randomly selected subsets of 8, 16, 32, 64, and 128 reports. RESULTS: We calculated the F1 evaluation metric weighted by support (number of true instances for each label) for each data field using the development test set. When working with the full training corpus (n=2066), convolutional networks perform the best (mean weighted F1 0.968 across all 12 clinical data elements). However, under smaller data conditions with less annotated data for training, boosting typically performs best. Moreover, with only 32 labeled reports we are able to achieve a mean weighted F1 score of 0.91 across fields. CONCLUSIONS: An NLP pipeline using both traditional statistical and machine-learning based methods can extract detailed prostate cancer pathology data from unstructured free text pathology reports with high accuracy, even with small sets of training data. Source of Funding: None San Francisco, CA; Berkeley, CA; San Francisco, CA; Berkeley, CA© 2019 by American Urological Association Education and Research, Inc.FiguresReferencesRelatedDetails Volume 201Issue Supplement 4April 2019Page: e1031-e1032 Advertisement Copyright & Permissions© 2019 by American Urological Association Education and Research, Inc.MetricsAuthor Information Anobel Odisho* More articles by this author Briton Park More articles by this author Nicholas Altieri More articles by this author William Murdoch More articles by this author Peter Carroll More articles by this author Matthew Coopberberg More articles by this author Bin Yu More articles by this author Expand All Advertisement PDF downloadLoading ...

Read the paper · More papers on PaperTik