Text Classification in USENET Newsgroups: A Progress Report
Scott Weiss, Simon Kasif, Eric Brill · 1996
We report on our investigations into topic classification with USENET newsgroups. Our framework is to determine the newsgroup that a new document should be posted to. We train our system by forming "metadocuments" that represent each topic. We discuss our experiments with this method, and provide evidence that choosing particular documents or words to use in these models degrades classification accuracy. We demonstrate SMART's deficiencies in retrieving documents from a newsgroup, and compare humans and SMART on a simple similarity task. We then describe a technique called classification-based retrieval for finding documents similar to a query document. 1 A Domain For Text Classification As the heralded "information superhighway" continues to expand, the need for useful methods of navigating it becomes greater. There are already a number of applications that aid the user. These include the WebWatcher at CMU[1], the LIRA system at Stanford[2], and the Webhunter and Letizia projects at ...