A probabilistic learning approach for document indexing

Norbert Fuhr, Chris Buckley · ACM Transactions on Information Systems · 1991

We describe a method for probabilistic document indexing using relevance feedback data that has been collected from a set of queries.Our approach is based on three new concepts: (1) Abstraction from specific terms and documents, which overcomes the restriction of limited relevance information for parameter estimation.(2) Flexibility of the representation, which allows the integrationof new text analysis and knowledge-based methods in our approach as well as the consideration of document structures or different types of terms.(3) Probabilistic learning or classification methods for the estimation of the indexing weights making better use of the available relevance information, Our approach can be applied under restrictions that hold for real applications.We give experimental results for five test collections which show improvements over other methods.

Read the paper · More papers on PaperTik