Automatic Generation of Background Text to Aid Classification.

Sarah Zelikovitz, Robert Häfner · 2004

We illustrate that Web searches can often be utilized to gen-erate background text for use with text classification. This is the case because there are frequently many pages on the World Wide Web that are relevant to particular text classifi-cation tasks. We show that an automatic method of creation of a secondary corpus of unlabeled but related documents can help decrease error rates in text categorization problems. Fur-thermore, if the test corpus is known, this related set of in-formation can be tailored to match the particular categoriza-tion problem in a transductive approach. Our system uses WHIRL, a tool that combines database functionalities with techniques from the information retrieval literature. When there is a limited number of training examples, or the process of obtaining training examples is expensive or difficult, this method can be especially useful.

Read the paper · More papers on PaperTik