Comparative evaluations of robust and accurate F0 estimates in reverberant environments

Masashi Unoki, Toshihiro Hosorogiya, Yuichi Ishimoto · IEEE International Conference on Acoustics Speech and Signal Processing · 2008

This paper reports comparative evaluations of the method we previously proposed of estimating fundamental frequency (F0) based on complex cepstrum analysis with nine typical methods over huge speech-sound datasets in both artificial and realistic reverberant environments (in room acoustics). They involve several classic algorithms (Cepstrum, AMDF, LPC, and modified autocorrelation) and a few modern algorithms (TEMPO, YIN, and PHIA). The comparative results revealed that the percentage correct rates of the estimated FOs using them were drastically reduced as the reverberation time increased while Foestimated with the proposed method was completely robust and accurate. They also demonstrated that homomorphic analysis and the concept of a source-filter model were relatively effective for estimating Fo. The results also demonstrated that it was much better than the previously reported methods in terms of robustness and providing accurate Foestimates in both artificial and realistic reverberant environments.

Read the paper · More papers on PaperTik