Clustering sequential data into hierarchical patterns
Xinying Song, Johnson Apacible · 2011
Sequential data, i.e. text string, is a common yet important data type. Automatically discovering patterns for sequential data is useful but challenging. In this paper, we address this task by clustering strings into hierarchical patterns. Such pattern hierarchy is particularly helpful for users to discover meaningful patterns as well as to interpret the encapsulated knowledge. We present the clustering algorithm in details and evaluate it on a large, real dataset of street addresses. The experiments demonstrate the effectiveness of our approach, making it a useful tool for analyzing and interpreting sequential data.