A Hybrid Method of Linguistic Features and Clustering Approach for Identifying Biomedical Named Entities
Ebtisaam Alharbi, Sabrina Binti Tiun · Asian Journal of Applied Sciences · 2015
Named entity is a term that has been widely used in the field of Natural Language Processing (NLP).It contains the names of persons, organizations, locations, dates and currencies.The process of extracting such names called Named Entity Recognition (NER).Biomedical Named Entity Recognition (BNER) is one of the fields that contains variety of named entities such as genes, DNA, RNA, chemical compounds.The key characteristic behind BNER lies on selecting an appropriate method that has the ability to identify the named entities effectively.Each entity (e.g., DNA, RNA, drugs) has its own features which are different from the others.Recently, identifying chemical compounds have caught researchers' attentions due to the various types of entities that included.Many approaches have been proposed in terms of extracting chemical compounds however, most of these approaches depend on a supervised learning techniques where the class label is predefined.In fact, chemical compounds have tremendous types of entities which requires more analytical categorization.For instance, Tramadol and Aspirin are drugs but each of them belongs to different classes of drugs.Hence, investing the unsupervised learning techniques may enrich these classifications.This study attempts to address the role of unsupervised learning specifically clustering K-means approach combining with feature extraction method including Part-of-Speech (POS) tagging and affixes (suffixes and prefixes).The experimental results of the proposed method have demonstrated an enhancement by obtaining 90% of F-measure.It is concluded that future efforts may concentrate on utilizing more linguistic features which could improve effectiveness.