Semantic Similarity for Music Retrieval

Luke Barrington, Douglas Turnbull, David Torres-Moreno, Gert R. G. Lanckriet · 2007

We present a query-by-example system for content-based music information retrieval by ranking items in a database based on semantic similarity, rather than acoustic similarity, to a query example. The retrieval system is based on semantic concept models that are learned from the CAL-500 data set containing both audio examples and their text captions. Using the concept models, the audio tracks are mapped into a semantic feature space, where each dimension indicates the strength of the semantic concept. Audio similarity and retrieval is then based on ranking the database tracks by their similarity to the query in the semantic space. 1 MODELING AUDIO AND SEMANTICS Our query-by-example music information retrieval (MIR) system takes an audio track as a query and retrieves new audio tracks that have similar semantic descriptions to the query track. For example, given a piece of music that a listener might describe as “crazy guitar rock with a screaming female singer that makes me want to get up and dance”, our system ranks all retrievable songs by how well they fit this description. The system is based on the models of [9, 3] which have shown promise in the domains of audio and image retrieval. Audio models are learned from a database of audio tracks with associated text captions that describe the audio content: D = {(A (1) , c (1)),..., (A (|D|) , c (|D|))} (1) where A (d) and c (d) represent the d-th audio track and the associated text caption, respectively. Each caption is a set of words from a fixed vocabulary, V. We train our system using the semantic labels from the CAL-500 data set [9] of 500 songs, each annotated by at least 3 humans using up to 200 words. We require that each word be positively associated with at least 10 songs, resulting in a vocabulary of 146 words (|V | = 146).

Read the paper · More papers on PaperTik