Clustering approaches to text categorization

Hiroya Takamura · Institutional Repositories DataBase (IRDB) · 2003

The aim of this thesis is to improve accuracy of text categorization, which is the basis for various applications such as e-mail classification and Web-page classification. Among the various possible approaches to this aim, two clustering approaches and an application of a new kernel (similarity function) are discussed in this thesis. Although clustering is usually regarded as an unsupervised learning method and categorization as a supervised learning, we show that clustering can be used to improve accuracy of text categorization. The first clustering approach proposed is co-clustering of words and texts. In a number of previous probabilistic approaches, texts in the same category are implicitly assumed to have an identical distribution over words. We empirically show that this assumption is not accurate, and propose a new framework based on a co-clustering technique to alleviate this problem. In this method, training texts are clustered so that the assumption is more likely to be true, and at the same time, features are also clustered in order to tackle the data sparseness problem. We succeeded in improving

Read the paper · More papers on PaperTik