Document classification for focused topics
Russell Power, Jay Chen, Trishank Karthik Kuppusamy, Lakshminarayanan Subramanian · 2010
Feature extraction is one of the fundamental challenges in im-proving the accuracy of document classification. While there has been a large body of research literature on document clas-sification, most existing approaches either do not have a high classification accuracy or require massive training sets. In this paper, we propose a simple feature extraction al-gorithm that can achieve high document classification ac-curacy in the context of development-centric topics. Our feature extraction algorithm exploits two distinct aspects in development-centric topics: (a) most of these topics tend to be very focused (unlike semantically hard classification top-ics such as chemistry or banks); (b) due to local language and cultural underpinnings in these topics, the authentic pages tend to use several region specific features. Our algorithm uses a combination of popularity and rarity as two separate metrics to extract features that describe a topic. Given a topic, our output feature set comprises of: (i) a list of popular key-words closely related to the topic; (ii) a list of rare keywords closely related to the topic. We show that a simple joint clas-sifier based on these two feature sets can achieve high classi-fication accuracy while each feature sub-set in itself is insuf-ficient. We have tested our algorithm across a wide range of development-centric topics. 1.