TREC 11 Experiments at NII: The Effects of Virtual Relevant Documents in Batch Filtering.

Kyung‐Soon Lee, Kyo Kageura, Akiko Aizawa · Text REtrieval Conference · 2002

Researches on document retrieval, text categorization and routing have shown the effects of learning by sampling relevant documents or non-relevant document from training set. Allan et al. (1995) considered only the top K non-relevant documents, which is the same number of all known relevant documents in the training set to learn a routing query. This is motivated by the need to have a balance between the number of the relevant and the negative documents in Rocchio’s learning. Singhal et al. (1997) selectively used the nonrelevant documents that belong to a query’s domain to learn the feedback query. Kwok and Grunfeld (1997) selected the best training subset of the relevant documents for creation of a feedback query based on genetic algorithm. Most sampling techniques in machine learning aim at the reducing the size of training set.

Read the paper · More papers on PaperTik