The TREC-4 Filtering Track.

David D. Lewis · 1995

The TREC-4 filtering track was an experiment in the evaluation of binary text classification systems. In contrast to ranking systems, binary text classification systems may need to produce result sets of any size, requiring that sampling be used to estimate their effectiveness. We present an effectiveness measure based on utility, and two sampling strategies (pooling and stratified sampling) for estimating utility of submitted sets. An evaluation of four sites was successfully carried out using this approach. 1 Introduction The goal of the TREC-4 filtering track was to develop methods for evaluating binary text classification systems, and try out those methods on real data. A secondary goal was to give TREC participants their first chance to evaluate approaches to binary text classification. This paper begins by defining binary text classification and presenting some applications of it. We then discuss a particular binary text classification task, filtering, used in TREC-4. The effect...

Read the paper · More papers on PaperTik