Data fusion of machine-learning methods for the TREC5 routing task (and other work)
Kwong Bor Ng, David Loewenstern, Chumki Basu, Haym Hirsh, Paul B. Kantor · Text REtrieval Conference · 1996
The goal of the document routing task is to extrapolate from documents judged relevant or irrelevant for each of a set of topics accurate procedures for assessing the relevance of future documents for each topic. Rather than viewing different approaches to this problem as winner-takes-all competitors, we view them as potentially complementar methods, each exploiting different sources of information. This paper describes two quite different machine-learning approaches to the document routing task, and two approaches to combining their results to perform relevance assessments on new documents. We also describe an approach to the confusion task based on n-grams that allow approximate matches