Clustering sequence data using hidden Markov model representation
Cen Li, Gautam Biswas · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1999
This paper proposes a clustering methodology, for sequence data, using hidden Markov model (HMM) representation. The proposed methodology improves upon existing HMM-based clustering methods in two ways: (i) it enables HMMs to dynamically change its model structure, to obtain a better fit model for data during the clustering process, and (ii) it provides objective criterion function, to select the optimal clustering partition. The algorithm is presented in terms of four nested levels of searches: (i) the search for the optimal number of clusters in a partition, (ii) the search for the optimal structure for a given partition, (iii) the search for the optimal HMM structure for each cluster, and (iv) the search for the optimal HMM parameters for each HMM. Preliminary results are given to support the proposed methodology.