Probabilistic Information Retrieval Based on Conceptual Overlap in Semantic Web Ontologies

Markus Holi, Eero Hyvönen · 2004

Abstract. Information retrieval systems have to deal with uncertain knowledge and query results should reflect this uncertainty in some manner. However, Se-mantic Web ontologies are based on crisp logic and do not provide well-defined means for expressing uncertainty. We present a new probabilistic method to ap-proach the problem. In our method, degrees of subsumption, i.e., overlap be-tween concepts can be modeled and computed efficiently using Bayesian net-works based on RDF(S) ontologies. Degrees of overlap indicate how well an in-dividual data item matches the query concept, which can be used as a well-defined measure of relevance in information retrieval tasks. 1 ONTOLOGIES AND INFORMATION RETRIEVAL A key reason for using ontologies in information retrieval systems, is that they enable the representation of background knowledge about a domain in a machine understand-able format. Humans use background knowledge heavily in information retrieval tasks [8]. For example, if a person is searching for documents about Europe she will use her background knowledge about European countries in the task. She will find a doc-

Read the paper · More papers on PaperTik