Unsupervised Segmentation of Audio Speech Using the Voting Experts Algorithm
Matthew Miller, Peter H. W. Wong, Alexander Stoytchev · 2009
Human beings have an apparently innate ability to segment continuous audio speech into words, and that ability is present in infants as young as 8 months old.This propensity towards audio segmentation seems to lay the groundwork for language learning.To artificially reproduce this ability would be both practically useful and theoretically enlightening.In this paper we propose an algorithm for the unsupervised segmentation of audio speech, based on the Voting Experts (VE) algorithm, which was originally designed to segment sequences of discrete tokens into categorical episodes.We demonstrate that our procedure is capable of inducing breaks with an accuracy substantially greater than chance, and suggest possible avenues of exploration to further increase the segmentation quality.