Graph-based Representations for Text Classification
Kjetil Valle · 2011
This paper presents a graph-based method for document representation intended for text classification with the vector space model. Terms are weighted by their centrality in networks constructed from the text. We evaluated a wide range of centrality measures, and experimented with two graph representations — co-occurrence networks and dependency networks. We compared the graph-based representations to the classical term frequency (TF) and term frequency-inverse document frequency (TF-IDF) based representations for classification. The graph-based representations performed better than the frequency-based measures on two datasets with different characteristics. We also found that representations considering only information local to each document, analogous of the TF measure, outperformed those including global information about the entire document collection similar to TF- IDF.