Incorporating Cumulative Mean Normalized Difference Function Towards Intepretable Monophonic Singing Voice Pitch Extraction
Xuefei Li, Chengcheng He · 2024
Pitch estimation plays an important role in various music processing and music information retrieval applications. The traditional methods for pitch estimation contain rich prior knowledge but struggle to accurately estimate pitch in noisy signals. In contrast, deep learning-based approaches perform well in noisy environments, but these methods have poor algorithmic interpretability. In this paper, we introduce the digital signal processing (DSP) process of the classical Yin algorithm into neural network design and propose an improved pitch estimation model for monophonic singing voice, namely YinNet. Specifically, we use neural networks to stimulate the cumulative mean normalized difference function (CMNDF) in Yin and design a Transformer-based post module. Experimental results show that our method outperforms CREPE, DeepF0 and other DSP-based methods significantly in both noisy scenarios and inter-corpus testings. In addition, the parameter size of our model is only 14% of CREPE, leading to a substantial reduction in computational complexity.