Data Augmentation for Layperson’s Medical Entity Linking Task

Annisa Maulida Ningtyas, Allan Hanbury, Florina Piroi, Linda Andersson · Forum for Information Retrieval Evaluation · 2021

Due to the vast amount of health-related data on social media, it is beneficial to monitor health-related issues experienced by the users, such as monitoring adverse drug effects. This problem is known as the Medical Entity Linking (MEL) task, which identifies informal medical terms and normalizes them to corresponding formal medical concepts (Medical Concept Normalization, or MCN). One of the challenges of MEL when working with layperson language is the scarcity of training data and the low coverage of laypersons’ health vocabulary. We apply data augmentation methods to existing data sets to increase the size of training corpora and expand the coverage of colloquial health vocabulary. We treat the informal medical phrase identification task as a sequence labelling task for Named Entity Recognition (NER). An extensive evaluation of NER shows that data augmentation may introduce noise that affects the performance of identifying informal medical phrases. Our further investigation on false positive in NER task finds that most of the false positive terms that occurred as a result of augmentation could be identified as medical entities and correctly normalized to SNOMED-CT medical concepts. In contrast, the augmentation techniques in MCN task can significantly improve performance on both the CADEC and PsyTAR data sets, particularly for paraphrase-based augmentation and combinations of augmentation techniques.

Read the paper · More papers on PaperTik