Optimizing Apache cTAKES for Disease/Disorder Template Filling: Team HITACHI in the ShARe/CLEF 2014 eHealth Evaluation Lab.
Nishikant Johri, Yoshiki Niwa, Veera Raghavendra Chikka · 2014
Abstract. This paper describes an information extraction system de-veloped by team Hitachi for “Disease/Disorder Template filling ” task or-ganized by ShARe/CLEF eHealth Evaluation Lab 2014. We approached the task by building a baseline system using Apache cTAKES. We sub-mitted two separate runs; in our first run, rule based assertion module predict the norm slot value of assertion attributes excluding training data knowledge. However assertion module is changed to machine learning-based in second run. We trained models for Course modifiers, Severity modifier and Body Location relation extractor and applied a variety of rule based post processing including structural parsing. We performed two layer search on UMLS dictionary for refinement of body location. Eventually, we created rules for temporal expression extraction and also used them as features for model training of DocTime. We followed a dictionary matching technique for cue slot value detection in Task 2b. Evaluation result of test data showed that our system performed very well in both subtasks. We achieved the highest accuracy 0.868 in norm value detection, strict F1-score 0.576 and relaxed F1-score 0.724 in cue slot value identification, indicating promising enhancement on baseline system.