Evaluating distributed word representations for capturing semantics of biomedical concepts

MUNEEB TH, Sunil Kumar Sahu, Ashish Prabhu Anand · 2015

Recently there is a surge in interest in learning vector representations of words using huge corpus in unsupervised manner.Such word vector representations, also known as word embedding, have been shown to improve the performance of machine learning models in several NLP tasks.However efficiency of such representation has not been systematically evaluated in biomedical domain.In this work our aim is to compare the performance of two state-of-the-art word embedding methods, namely word2vec and GloVe on a basic task of reflecting semantic similarity and relatedness of biomedical concepts.For this, vector representations of all unique words in the corpus of more than 1 million full-length research articles in biomedical domain are obtained from the two methods.These word vectors are evaluated for their ability to reflect semantic similarity and semantic relatedness of word-pairs in a benchmark data set of manually curated semantic similar and related words available at http:// rxinformatics.umn.edu.We observe that parameters of these models do affect their ability to capture lexicosemantic properties and word2vec with particular language modeling seems to perform better than others.

Read the paper · More papers on PaperTik