An aspect based document representation for event clustering
Wim De Smet, Marie‐Francine Moens · Lirias · 2009
We have studied several techniques for creating and comparing content represen-tations of textual documents in the field of event detection. We define a document as a collection of aspects, i.e. disjoint components that reveal (latent) topics and/or extracted information such as named entities. As underlying models we consider the vector space model and probabilistic topic models based on Latent Dirichlet Allocation. We also investigate the value of dependencies between the aspects, which are reflected by importance factors. We apply and evaluate our techniques on event detection in Wikinews, where we cluster news stories that discuss the same event. We found that the split representations yield the best event detection results compared to the ground-truth event clusters. Our methods for aspect detec-tion, for learning the importance factors of the aspects, and for event clustering are completely unsupervised. 1