Audio Speech Segmentation Without Language-Specific Knowledge
Kevin Gold, Brian Scassellati · eScholarship (California Digital Library) · 2006
Speech segmentation is the problem of finding word boundaries in spoken language when the underlying vocabulary is still unknown.Here we show that a system with no phonemic knowledge can find word boundaries.The system first subdivides an utterance by recursively clustering similar parts of the signal together until the cepstral coefficient variance is low within each new segment.These segments are then used as inputs to a perceptron-like algorithm that finds repeated segments across utterances.With only a few sample utterances, and no previous linguistic knowledge, the system can find the words that were repeated across utterances and identify new utterances that contain those words.The findings show that the assumption of a phoneme classification module is not necessary for a "minimum description length" (Brent & Cartwright, 1996;de Marcken, 1996) explanation of word segmentation.