RAVE Reviews: Acquiring relevance assessments from multiple users
K. Belew, John Hatton · 1996
As the use of machine learning techniques in IR increases, the need for a sound empirical methodology for collecting ud assessing users’ opinions- "relevance feedback "- becomes critical to the evaluation of system performance. In IR the typical assessment procedure relies upon the opinion of a single individual, an "expert " in the corpus ’ domain of discourse. Apart from the logistical dit~ulties of gethefieg multiple opinions, whether any one, "omnlmcent " individual is capable of providing reliable a.t. about the appm~ate set of documents to be reUievod remains a foundational issue within IR. "l~is paper responds to such critiques with a new methodology for collecting relevance assessments that combines evidence from mulitple human judges. RAVe is a suite of software routines that allow an IR experlmeuter to effectively collect large numbers of relevance assessments for an arb/trary document corpus. paper sketches our assumptions about the cognitive activity of the providing relvance assessments, and the design issues involved in: identifying the documents o be eval~3_ed; allocating subjects ’ time to provide the most infommtive assessments; and aggregating multiple users ’ opinions into a binary predicate of "relevut. " Finally, we present lnel/ml-my a,t, gathen ~ by RAVe from subjects.