Representation models for text classification
George Giannakopoulos, Petra Mavridi, Γεώργιος Παλιούρας, George Papadakis, Konstantinos Tserpes · 2012
Text classification constitutes a popular task in Web research with various applications that range from spam filtering to sentiment analysis. To address it, patterns of co-occurring words or characters are typically extracted from the textual content of Web documents. However, not all documents are of the same quality; for example, the curated content of news articles usually entails lower levels of noise than the user-generated content of the blog posts and the other Social Media.