Towards Computational Auditory Scene Analysis: Melody Extraction from Polyphonic Music

Karin Dressler · 2012

Abstract. This paper describes an efficient method for the identifica-tion of the melody voice from the frame-wise updated magnitude and frequency values of tone objects. Most state of the art algorithms em-ploy a probabilistic framework to find the best succession of melody tones. Often such methods fail, if there are several musical voices with a comparable strength in the audio mixture. In this paper, we present a computational method for auditory stream segregation that processes a variable number of simultaneous voices. Although no statistical model is implemented, probabilistic relationships that can be observed in melody tone sequences are exploited. The method is a further development of an algorithm which was successfully evaluated as part of a melody ex-traction system. While the current version does not improve the overall accuracy for some melody extraction data sets, it shows a superior per-formance for audio examples which have been assembled to show the effects of auditory streaming in human perception.

Read the paper · More papers on PaperTik