Using Proximity Operators for Document Representation in Web Search.

Jesus Serrano‐Guerrero, José Á. Olivas, Javier de la Mata · International Conference on Artificial Intelligence · 2006

Abstract It is presented a different way for representing documents in information retrieval systems. The relation degree between the index term pairs which describe the document is proposed as the main factor to represent a document. This document representation approach is based on two kinds of different relations between index terms: the physical relation and the semantic relation. Keywords: proximity operators, search engine, semantic relations, document representation. 1. Introduction Classic models (Boolean, probabilistic and vector space) in information retrieval are based on the document representation by means of a set of index terms. Usually, to calculate the similarity degree between a document and a query, the main factor to keep in mind is the frequency of the index terms of the document. This fact causes the lack of conceptual relevance of the query results. The user query is also limited because the user can only use terms to express it, the query is a “lexical focused” query and not a “conceptually focused” query . The classic systems only use the terms appeared on the documents as index terms, but there are some different works that use ontologies in the indexing process. For example, the semantic indexing technique [6] proposed by Mihalcea uses WordNet [9] to disambiguate the polisemic words. This method improves the results introducing new related terms in the document representation and in the query expansion stage. Other strategy is to use electronic dictionaries to index the documents [3].

Read the paper · More papers on PaperTik