A MeSH term based distance measure for document retrieval and labeling assistance

Jörg Ontrup, Tim W. Nattkemper, Olaf Gerstung, Helge Joachim Ritter · 2004

For biomedical and pharmaceutical research, the PUBMED database of the NLM (National Library of Medicine) has become a viable platform. It provides the means for profound investigations of past and related research in daily scientific work. One basic aspect is the search for articles related to a certain research topic. In order to express relatedness many text-mining or document retrieval approaches make use of the "bag of words" model in which unstructured text is represented as a vector of word counts. Since full length articles are not commonly available, many systems generate feature vectors from abstract data only - therefore limiting the explanatory power of their feature space. Since MeSH (Medical Subject Headings) assigned by human experts cover full length articles, we propose for the first time a nonEuclidean document distance measure based on MeSH tree structures. We quantitatively evaluate the approach in comparison to a standard vector space approach and a hybrid version of both. The MeSH-based showed promising results, yet it is still surpassed by the vector space model.

Read the paper · More papers on PaperTik