An evaluation of term dependence models in information retrieval
Gerard Salton, Chris Buckley, C. Yu · 1982
In practical retrieval environments the assumption is normally made that the terms assigned to the documents of a collection occur independently of each other. The term independence assnmption is unrealistic in many cases, but its use leads to a simple retrieval algorithm. More realistic retrieval systems take into account dependencies between certain term pairs and possibly between term triples. In this study, methods are outlined for generating dependency factors for term pairs and term triples and for using them in retrieval. Evaluation output is included to demonstrate the effectiveness of the suggested methodologies. i. Term Dependency Models From a decision-theoretic viewpoint, the information retrieval task is con-trolled by two probabilistic parameters which specify for each document of a collec-tion the probability of relevance, and the probability of nonrelevance, with respect to a particular query. The larger the probability of relevance, and the smaller the probability of nonrelevance, the greater is the retrieval probability for the given item. Consider in particular an item ~ in the data base represented by binary attri-butes (Xl,X2,...,Xn), where x i takes on the values i or 0 depending on whether the ith attribute is or is not assigned to item ~. For each item ~ and each query Q, it is in principle possible to generate the two parameters P(xJrel) and P(x[nonrel), representing the probabilities that a relevant and a nonrelevant item, respectively, has vector representation ~. Using decision theoretic considerations, it is easy to show that an optimal retrieval rule will rank the documents in decreasing order according to the expression P(x!ret) P(~[nonrel) (I) That is, given two items x and v, x should be retrieved ahead of ~ whenever the value of expression (I) for x exceeds the corresponding value for ~. [1-5]