Automatic scoring of sung melodies in comparison with human performance
Masuzo Yanagida, Yumiko Mizuno, Tomoya Matsunaga · 2010
Performance of transcribing acoustic signals into music notation is compared between an automatic transcribing system recently developed by the authors’ group and human scorers, focusing mainly on tone height using two kinds of singing sounds : one, sung by human singers and the other, synthesized with a commercially sold singer system. As test melodies should be unknown to human scorers, several melodies were designed for this particular experiment. Those melodies are divided into two sets of melody groups: one for human singing and the other for synthetic singing. The reason why the test melodies were not made common for human singing and synthetic singing is that each test molody should be unknown to human scorers. Each set is sub-classified into three tonality levels, highly tonal, less tonal and atonal. Significant difference is recognized in correct tone height transcription rate among three tonality levels both for human scorers and the automatic transcription system for melodies sung by human singers. Although results obtained for melodies sung by human singers are somewhat vague as the correct answers are hard to be defined, but results obtained for melodies sung by synthetic voice were just as expected that transcription performance of the system is far superior to that of human scorers for atonal melodies or tone series, and significantly superior for low tonality melodies, but no significant difference for high tonality melodies .