Piano Multi-Pitch Estimator Using CNN-Stacked LSTM
Christhopher Ravian Hartono, Salim Hartono, Benyamin Budiharja, Ivan Sebastian Edbert, Derwin Suhartono · 2022
This research explores several variations of the automatic music transcription method, specifically in the pitch estimation task. Pitch estimation in this research mainly converts an acoustic piano song recording into a digitally transcribed song format. First, several techniques, including short-time Fast-Fourier transform and constant-Q transform, provide a spectrogram representation of a wav piano recording. Then it is fed into a combination of Convolutional Neural Network (ConvNet) and Long Short-Term Memory (LSTM) neural network. This transcription is a digitally transcribed song in the format of a MIDI file. For training purposes, the MAESTRO dataset was used for conducting training, which every training varies the learning rate value and spectrogram representation. This research obtains a model that can perform pitch estimation with a 90.14% F1 score and an average user evaluation of 8.4 out of 10.