Text Categorization with All Substring Features

Daisuke Okanohara, Jun’ichi Tsujii · 2009

This paper presents a novel document classification method using all substrings as features. Although tokenized words are not enough for determining a class of a document, learning by using all substrings has a prohibitive computational cost because the number of all candidate substrings can be very large. We show that the idea of equivalent classes of substrings can help determine all effective substrings exhaustively in linear time. Moreover, by applying L1 regularization to our model, we obtain a compact result, which makes an inference extremely efficient in time and space, and robust even if we use substrings of all lengths. In experiments, we show that our method can extract effective substrings efficiently, and achieved more accurate results and the its inference was faster than the results using previous methods.

Read the paper · More papers on PaperTik