Mining for Surprise Events Within Text Streams
Paul Whitney, Dave W. Engel, Nick Cramer · 2009
Text streams are a fundamental source of information that can be used to detect and characterize strategic intent of individuals and organizations as well as for detecting abrupt or surprising events within communities. In this paper we describe our algorithm development and analysis methodology for mining the evolving content in text streams. Text streams include news, press releases from organizations, speeches, Internet blogs, etc. Specifically, an analyst may need to know if and when the topic within a text stream changes. Much of the current text-feature methodology is focused on understanding and analyzing a single static collection of text documents. Corresponding analytic activities include summarizing the contents of the collection, grouping the documents based on similarity of content, and calculating concise summaries of the resulting groups. The approach reported here focuses on taking advantage of the temporal characteristics in a text stream to identify relevant features (such as change in content), and also on the analysis and algorithmic methodology to communicate these characteristics to a user. We present a variety of algorithms for detecting essential features within a text stream. Our approach for communicating the information back to the user is to identify feature (word/phrase) groups. These resulting algorithms form the basis of developing software tools for a user to analyze and understand the content of text streams. We present analysis results using both news information and abstracts from technical articles, and show how these algorithms provide understanding of the contents of these text streams. A critical finding is that the characteristics we used to identify features in a text stream are uncorrelated with the characteristics used to identify features in a static document collection.