When Less is More: Relevance Feedback Falls Short and Term Expansion Succeeds at HARD 2005
Fernando G. Diaz, James Allan · 2006
We used clarification forms to study passage term feedback. When compared against pseudo-relevance feedback with an extremely large external corpus, we found that passage feedback resulted in a reduction in performance while term feedback significantly improved recall. 1. OVERVIEW UMass tested several new techniques in the HARD track this year. First, we developed a new baseline pseudo-relevance feedback technique based on expanding from a large, external corpus [6]. Our results indicate that externally expanding a query can result in improvements over passage feedback. Second, we present a new feedback technique which incorporates both relevance and non-relevance information. Our results indicate that this regularization-based technique can improve performance when used in conjunction with traditional feedback techniques. Third, we present two successful term feedback techniques. Our most successful technique exploits a structured query language in order to significantly improve recall. 2. BASELINE ALGORITHMS We used the Indri retrieval engine for retrieval experiments [7]. We did not submit any runs which did not perform pseudo-relevance feedback. We experimented with a mixture of relevance models for our baseline ranking algorithm (also introduced in our Robust runs this year) [6]. Readers should consult the refered work for a more thorough description of parameters. Mixture parameters were set to P (aquaint) = 1 for MASSbaseTRM3 and P (bignews) = 1 for MASSbaseTEE3. No dependence models were used in our baselines. 3. CLARIFICATION FORMS Our clarification form consisted of several pages of dialog with the searcher. This dialog followed three phases. In the first phase, the searcher was presented with passages from which to judge document relevance. The second phase consisted of term-based feedback; searchers were asked to judge the expected frequency of terms in relevant documents. 3.1 Passage Presentation In previous years, we found that passages acted as suitable surrogates for documents when being judged for relevance [1]. We divided each of the top 5 documents in the MASSbaseEE3 run into 150-word, half-overlapping passages. We