Query-by-singing/humming systems based on speech recognition technologies

Motoyuki Suzuki, Akinori Ito, Shozo Makino · The Journal of the Acoustical Society of America · 2007

While many query-by-singing/humming systems have been developed, several problems have prevented their practical use. For example, pitch extraction methods are commonly used in QbSH systems, but they often yield inaccurate information. To rectify this problem, a QbH system (i.e., no lyrics required), which uses multiple pitch candidates, has been developed. For each input note, multiple pitch candidates are extracted, and all combinations of pitch candidates are considered in the retrieval step. This system has shown higher performance than QbH systems, which use single pitch candidates. QbS systems (i.e., lyrics required) have also been developed. In this case, one of the biggest problems is to segment an input singing voice into musical notes. Several methods have been proposed. However, it is hard to segment the singing voice without linguistic information. In the proposed system, a lyrics recognizer outputs several hypotheses for a recognition result. Each hypothesis gives time alignment information between the singing voice and the recognized text. Once this information is available, it becomes easy to segment the input singing voice into musical notes, and a hypothesis can be verified from a melodic point of view. This system, which has 156 songs in the database, yields 93% retrieval accuracy for a query consisting of five notes.

Read the paper · More papers on PaperTik