Hidden Markov Model Clustering of Acoustic Data
Matthew Butler · 2003
This dissertation explores methods for cluster analysis of acoustic data. Techniques developed are applied primarily to whale song, but the task is treated in as general a manner as possible. Three algorithms are presented, all built around hidden Markov models, respectively implementing partitional, agglomerative, and divisive clustering. Topology optimization through Bayesian model selection is explored, addressing the issues of the number of clusters present and the model complexity required to model each cluster, but available methods are found to be unreliable for complex data. A number of feature extraction procedures are examined, and their relative merits compared for various types of data. Overall, hierarchical HMM clustering is found to be an effective tool for unsupervised learning of sound patterns.