Enhancing Piano Transcription by Dilated Convolution
Xian Hui Wang, Lingqiao Liu, Qinfeng Shi · 2020
When detecting music pitch with deep learning, the most challenging issue is how to aggregate the multi-scale frequency information scattered in the spectrogram of a music signal. Traditionally, this aggregation was achieved via strided pooling and fully connected operations, or domain-specific sparse convolutions. In this paper we propose to use more general-purpose dilated convolution to this end. In particular, first we design a dilated convolutional acoustic model for piano pitch detection. This model leads existing acoustic models by large margins. Based on this model, we then design an automatic music transcription (AMT) system. When trained and tested on the MAPS dataset, this system outperforms existing AMT systems again by large margins. When trained and tested on the MAESTRO dataset, this system performs equally well as a state- of-the-art system, but has much better generalization capability.