Finding Hotspots in Document Collection
Wei Fu Peng, Chris H. Q. Ding, Tao Li, Tong Sun · 2007
Given a document collection, it is often desirable to find the core subset of documents focusing on a specific topic. We propose a new algorithm for this task. Document clustering aims at par- titioning the document-term datasets into differ- ent groups by optimizing certain objective func- tions. However, they are not suitable for finding hotspots that are described by a small set of doc- uments with few tightly coupled terms. In this pa- per we propose a novel hotspot finding algorithm, DCC (Dense Concept Clustering) in document collections. DCC can extract distinct small top- ics with most representative documents and words simultaneously. The hotspots are dense bicliques in binary document-word matrices and they can be discovered sequentially one at a time using the generalized Motzkin-Straus formalism. The rep- resentative documents and words are tightly cor- related for concept descriptions. Experiments on real document datasets show the effectiveness of the proposed algorithm.