Time-domain Polyphonic Transcription using Self-generating Databases
Juan Pablo-Bello, Laurent Daudet, Mark B. Sandler · Journal of the Audio Engineering Society · 2002
We describe a new method for the estimation of multiple pitch information in recorded piano music. The method works in the time-domain and makes use of a self-generating database of all possible notes. First, we show how accurate polyphonic pitch detection can be achieved given an adequate database. Then an algorithm is proposed that generates the database from the music, using estimation of predominant pitches in the frequency-domain and pitch-shifting techniques. Both systems generate a MIDI representation of the original signal. This method -that can be generalized to any solo instrumentovercomes the usual constraints of the traditional frequency-domain approach regarding intervals and quantity of notes. INTRODUCTION AND OBJECTIVES Extracting meaning from music is a process that comes natural to human perception. There are very different levels of information, subjective (e.g. style, mood...) and objective (e.g. tempo, notes...), that can be extracted from music. Usually objective feature recognition hugely depends on the levels of training and knowledge that the listener possesses. The music transcription task (the process of converting audio to score) is an example of this: the more complex the musical input becomes, the more acute the need for prior knowledge. Automatic music transcription tries to recreate this task with computer algorithms, but becomes increasingly complicated when dealing with polyphonic music, that presents multiplicity of pitches and possibly timbres. Most monophonic transcription techniques are not applicable to this case, forcing researchers to switch views and to propose novel ways of tackling different aspects of this problem. Previous systems rely, almost as a rule, on the analysis of information in the frequency domain. In time-frequency repBELLO ET AL. TIME-DOMAIN POLYPHONIC TRANSCRIPTION resentations, periodicities in time, such as pitch, are represented as energy maxima in the frequency-domain. This suggests that the grouping of certain series of these energy maxima, generates patterns or structures that may be related to notes in a music file. Some proposed polyphonic transcription systems include predominant pitch estimation in separate frequency bands [1], data-fusion architectures for the grouping of supportive frequency-domain evidence [2, 3], statistical frameworks for the generation of frequency-events probability matrices [4, 5], peak location comparisons of harmonic spectra [6], etc. However, relying on the analysis of the frequencydomain data has some disadvantages. We find particularly relevant those caused by the harmonicity and the polyphony of the analysis signal: Harmonicity: Two sounds are considered to be harmonic, when they have a fundamental frequency ratio a : b (where a and b are positive integers). This implies that every b-th partial of Sound 1 overlaps every a-th partial of Sound 2 [7]. Intervals of the western musical scale commonly produce perfect or near-perfect harmonic relations (e.g. the fifth 2 : 3, the third 3 : 4). Hence simultaneous notes often generate important overlapping between partials (see fig. 1), thereby complicating the process of identifying notes. A common interval, the octave (1 : 2), is particularly difficult having all the partials of Sound 1 overlap all the partials of Sound 2. Proposed systems, relying on partials’ information, have trouble dealing with such intervals. Polyphony: When more than one note is present, peaks related to different pitches can lie in the same frequency bin and therefore are not uniquely identifiable. This problem gets progressively worse as more notes are added. Increasing polyphony often means increasing error rate for the abovementioned approaches. The objective of this paper is primarily to address polyphonic music transcription while avoiding, at least partially, the usual paradigm of analysis in the frequency domain. Secondly, to propose a system that extracts knowledge about the signal, from the signal itself. Here, our definition of “transcription” is limited to the estimation of onset times, durations and pitches of the notes being played. TRANSCRIPTION IN THE TIME DOMAIN Here, we propose a hybrid method, where the classical frequency-domain approach is improved by a time-domain recognition process. This enables a refinement of our results by taking into account the information contained in phase relationships, that are lost when only spectrum of sounds are analyzed. As we have seen in the last section, most errors (harmonicity and polyphony) occur when analyzing complex chords. In this section, let us present the methodology, assuming that we are restricted to the analysis of a short frame containing a single chord. The linear additive model As a first approximation, let us assume a linear additive model, where the waveform of a given note is independent (up to a global scaling) of the loudness; as well as of the presence of other notes. Let D = {xi}i=1...88 be the database containing the normalized waveforms (in the time domain) of the 88 individual piano notes. Under these assumptions, a chord will simply be a weighted linear sum of the individual Fig. 1: The problem of harmonicity with chords made of two notes related by simple intervals. Dark areas indicate frequency regions where partials overlap between the two notes. Top: third interval Middle: fifth interval Bottom: octave.