Research on Biological Literature Information Retrieval Based on Cluster Language Model
Wen Jian · Zhongwen xinxi xuebao · 2008
Recent researches present topic language model improves the performance of information retrieval,but many problems still has not been solved include data sparseness problem,synonymy and polysemy problems,smoo-thing the seen term or not seen term.All the problems are important to IR,especially in domain literature IR,for example biological literatures.In this paper,a new topic language model based on cluster was proposed.The work mainly included two aspects.First,documents were represented by concepts of ontology,and concept-based clustering is done using Fuzzy C-Means,the clustering result was considered as the topics of document collections.The probability of a document generating topics is estimated by the similarity between the document and each cluster.Then,the probability of topic generating words is estimated using Expectation Maximization algorithm.At last,Through integrating the above algorithms into the aspect model,our topic language model was formed.This new language model accurately describes the distributed probability of words in different topics and the probability of a document generating a topic.Moreover,it can partly solve synonymy and polysemy problems.The new method was evaluated on TREC 2004/05 Genomics Track collections.Experiments have shown that the retrieval performance has been improved by the new method compared with simple language model.