Mining Domain Information from Social Contents Based on News Categories

Yin‐Fu Huang, Chen-Ting Huang · 2014

In this paper, to help users find the domain information from tweets they are interested in and further support the work of domain explorers and social content retrieval, a classification method on the social contents in Twitter is proposed, based on news categories. The classification method uses traditional "bag of words" features and explicit features extracted from tweets to facilitate the classification. Since data sparseness is always a serious problem when considering "bag of words" features in the text classification, we further employ dimensionality reduction methods on "bag of words" features and observe their performances. The experimental results show that 1) our proposed method can achieve good performances in the tweet classification and 2) using dimensionality reduction methods can achieve higher accuracy than not using them.

Read the paper · More papers on PaperTik