Effectiveness of Weighted Searching in an Operational IR Environment.

Hanspeter Frei, Yonggang Qiu · 1993

this paper. First of all, Cirt is a front-end system connected to an operational IR host with real users, queries, databases, hosts, networks, etc. The system used for the research described in this paper was designed similarly and even the same host was used, namely Data-Star. Secondly, the results reported from the Cirt experiment were not really enthusiastic. This is in sharp contrast to our experiment that showed a significant improvement in weighted retrieval over Boolean retrieval. Although the authors claim in [Rob 90] that 'Weighted retrieval is capable of achieving results comparable to those obtained with Boolean searching,' the attentive reader notices that the authors actually expected quite a bit more. The main reason for the mediocre success of Cirt was the inability of the host to execute weighted queries. Rather, weighted queries had to be converted into sequences of Boolean queries that were sent to the host and evaluated as if they were issued by a normal Data-Star user. We experienced the same disappointment when we did a similar experiment a year earlier. It seems to be difficult to simulate weighted queries in a Boolean environment especially when some of the basic data like term frequency is not available. In addition, it is well-known that one of the strengths of weighted retrieval is caused by the usually large number of search terms employed. When using Cirt, the users used only a few - 3 search terms because of the time it took to process queries. Therefore, the actual potential of the weighted technique could never really surface. We profited from the fact that a real weighted retrieval algorithm has been built directly into the Data-Star system in the meantime. This algorithm accepts a weighted query q represented by a vector q = (a 1 , a 2...

Read the paper · More papers on PaperTik