Support Tensor Machines for Text Categorization

Deng Cai, Xiaofei He, Ji-Rong Wen, Jiawei Han, Wei‐Ying Ma · 2006

We consider the problem of text representation and categorization. Conventionally, a text document is represented by a vector in high dimensional space. Some learning algorithms are then applied in such a vector space for text categorization. Particularly, Support Vector Machine (SVM) has received a lot of attentions due to its effectiveness. In this paper, we propose a new classification algorithm called Support Tensor Machine (STM). STM uses Tensor Space Model to represent documents. It considers a document as the second order tensor in Rn1⊗Rn2, where Rn1 and Rn2 are two vector spaces. With tensor representation, the number of parameters estimated by STM is much less than the number of parameters estimated by SVM. Therefore, our algorithm is especially suitable for small sample cases. We compared our proposed algorithm with SVM for text categorization on two standard databases. Experimental results show the effectiveness of our algorithm. 1

Read the paper · More papers on PaperTik