Probabilistic dependence and logistic inference in information retrieval

Fredric C. Gey · University of California at Berkeley eBooks · 1993

This research presents a new model for probabilistic text and document retrieval, a model which is believed to be both theoretically sound and implementable. The model utilizes the technique of logistic regression to obtain equations which rank documents by probability of relevance as a function of document and query properties. Since this model infers probability of relevance from statistical clues present in the texts of documents and queries, we call it logistic inference. By transforming the distribution of each statistical clue into its standardized distribution (one with mean $\mu$ = 0 and standard deviation $\sigma$ = 1), the method allows one to apply logistic coefficients derived from a training collection to other document collections, with no loss of predictive power. The model is applied to three well-known information retrieval test collections, and the results are compared directly to the particular vector space model of retrieval which uses term-frequency/inverse-document-frequency (tfidf) weighting and the cosine similarity measure. In the comparison, the logistic inference method performs significantly better than (in two collections) or equally well as (in the third collection) the tfidf/cosine vector space model. The differences in performances of the two models were subjected to statistical tests to see if the differences are statistically significant or could have occurred by chance. This research also presents a study of probabilistic dependencies in the information retrieval situation. The CACM information retrieval test collection is examined to determine whether a null hypothesis of the binary independence model of probabilistic querying is valid. Statistical tests on query term pairs seem to lead us to reject this hypothesis.

Read the paper · More papers on PaperTik