On human capability and acoustic cues for discriminating singing and speaking voices

Yasunori Ohishi, Masataka Goto, Katunobu Itou, Kazuya Takeda · 2006

In this paper, acoustic cues and human capability for discriminating singing and speaking voices are discussed to develop an automatic discrimination system for singing and speaking voices. Based on the results of preliminary subjective experiments, listeners discriminate between singing and speaking voices with 70.0% accuracy for 200ms signals and 99.7% for one-second signals. Since even short stimuli of 200 ms can be correctly discriminated, not only temporal characteristics but also short-time spectral features can be cues for discrimination. To examine how listeners distinguish between these two voices, we conducted subjective experiments with singing and speaking voice stimuli whose voice quality and prosody were systematically distorted by using signal processing techniques. The experimental results suggest that spectral and prosodic cues complementarily contributed to perceptual judgments. Furthermore, a software system that can automatically discriminate between singing and speaking voices and such performances is also reported.

Read the paper · More papers on PaperTik