Automatic Topic Extraction from Temporal Text Streams
신용욱 · Seoul National University Open Repository (Seoul National University) · 2012
As social media services such as Twitter and Facebook are gaining popularity, the amount of information from these services is explosively growing. Most of them use temporal text stream to facilitate distribution of a huge volume of content they publish. Dynamic text stream is defined as a temporarily ordered set of documents published by a web site over time in order to facilitate syndication of content. In this context, many users subscribe to the streams to acquire up-to-date information through information aggregation and sharing services, and real-time search engines also increasingly utilize the streams to promptly find recent web content when it is produced. Compared to web pages, temporal text stream is a time-varying document since it continually publishes entries on some specific topics. In addition, it is a structured document that consists of several data elements such as title and description. In this thesis, we investigate a problem of extracting topics from the streams, considering the temporal and structural characteristics of temporal text stream. Specifically, the first part of the thesis considers a problem of identifying a feature set created from data elements constituting temporal text stream, with the aim of improving effectiveness of extracting persistent topics over the stream while at the same time reducing computational cost. With structural nature of the stream, it is necessary to investigate which data elements need to be selected to define a feature set for topic extraction. Furthermore, the temporal characteristic of the stream raises a problem of determining how many entries need to be considered for topic extraction. The second part of the thesis addresses a problem of detecting topics of persistent interests from a temporal text stream over time, considering its temporal characteristics. After defining three unique properties of the persistent topics, a graph-based topic extraction model using scoring functions to measure the properties is proposed. Finally, a novel automatic tagging model to detect informative terms from short entries of the temporal text stream is proposed following the framework of supervised approach. Traditional frequency-based term features are redefined so that they can address the properties of the entries created under the length limitation, and sequential dependencies between successive terms in an entry is considered. In addition, the proposed automatic tagging approach incorporates behavioral patterns by which users put informative terms into their entries. The thesis concludes with a discussion on potential extensions of this work.