Transcription of Audio to MIDI Using Deep Learning

Patrick J. Donnelly, Victoria Ebert · 2022

We investigate automatic transcription of polyphonic piano music with recurrent neural networks directly from raw audio. Previous approaches to this task rely on spectrograms or other features extracted for training, which requires timely pre-processing. In our approach, we train directly on sequences of 64 milliseconds of samples from the waveform. We evaluate our models on the MAESTRO dataset and achieve promising levels of pitch and onset accuracy but difficulty detecting the correct offset of the notes. Our approach also enables learning the note pitches and velocities at the same time. Together these results demonstrate the potential of relatively simple deep learning architectures to learn music transcription directly from raw audio.

Read the paper · More papers on PaperTik