Fast Entity Recognition in Biomedical Text
Amy Siu, Dat Ba Nguyen, Gerhard Weikum · MPG.PuRe (Max Planck Society) · 2013
In biomedical text mining, entity recognition is often an early task in the pipeline of analyzing free text.MetaMap, the de facto standard software tool for this task, employs much Natural Language Processing (NLP) machinery to recognize entities in UMLS (Unified Medical Language System), the largest metathesaurus.Knowing that the NLP machinery is time-consuming, and that UMLS is rich in lexical variations, this work investigates whether a fast, string similarity-based method can achieve results comparable to those of MetaMap.We implemented an NLP-light method that performs fast MinHash lookups via character trigram features.Starting with UMLS as the dictionary of entities, we select a subset whose entity names are short and thus amenable to a string similarity-based approach.We applied the method to both scientific literature and layman-oriented texts from Internet health portals.Our proposed method achieved up to 83% precision and 78% coverage under a strict rating scheme that penalizes failure in Word Sense Disambiguation (WSD), at a throughput of 1,720 PubMed abstracts or 175 web pages per minute.When compared to MetaMap, our proposed method achieved comparable precision and 13% less coverage using less than 1% of the time.