Topic Characterization of Full Length Texts Using Direct and Indirect Term Evidence

David E. Fisher · 1994

This project evaluates two families of algorithms that can be used to automatically classify general texts within a set of conceptual categories. The #rst family uses indirect evidence in the form of term#category co-occurrence data. The second uses direct evidence based on the senses of the terms, where a term's senses are designated by the categories that it is a member of in a thesaurus. The direct evidence algorithms incorporate varying degrees of indirect evidence as well. For these experiments a set of 3864 conceptual categories were derived from the noun hierarchyofWordNet, an on-line thesaurus. The co-occurrence data for the associational and disambiguation algorithms was collected from a corpus of 3,711 AP newswire articles, comprising approximately 1.7 million words of text. Each of the algorithms was applied to all of the articles in the AP corpus, with their performance evaluated both qualitatively and quantitatively. The results of these experiments show that both classe...

Read the paper · More papers on PaperTik