A Pattern Mining Approach to Study a Collection of Dutch Folk-Songs
Conklin, Darragh · Arrow - TU Dublin (Technological University Dublin) · 2016
ion level of the viewpoints should be high enough to capture variability in the melodies as caused both by the process of oral transmission and by variations in choices that were made in the process of transcription into music notation. To achieve a suitable level of abstraction, we measure relative values for all viewpoints derived from pitch or duration. For the current study we define the following viewpoints: phrpos, which records whether the note is the first in a phrase, the last in a phrase, or inside a phrase; intref, which represents the scale degree of the note given the key of the song; c3i(level), which records whether the metric level of a note is higher, lower or equal with respect to the previous note; c3(dur), which records whether the note is shorter, equal, or longer in duration than the previous note; c3(pitch), which records whether the note is higher, equal, or lower in pitch than the previous note; c5(pitch, 3), which records whether the note was approached by a leap (three semitones or larger), a step (smaller than a three semitones), or a unison, with distinction between ascending and descending intervals; and c5(pitch, 7), which records whether the note was approached by a leap (seven semitones or larger), a step (smaller than seven semitones), or a unison, with distinction between ascending and descending intervals. A feature is a tuple τ : v comprised of a viewpoint name τ paired with a value v. A feature set is a set of features, for example the feature set { c3(pitch) : − intref : M2 } contains two features, expressing that the pitch of the corresponding note is lower than that of the previous note, and is the major second (M2) of the scale. An event instantiates a feature set if all features in the set are true for the event. A feature set pattern is a sequence of feature sets, and a song instantiates a pattern (or, stated equivalently, the pattern occurs in the song) if the successive feature sets of the pattern instantiate successive events in the song in at least one place. For example, the patterns shown in Figure 1 have four feature sets, with different features in each of them. Following the method presented by Conklin (2010), a one vs. all strategy (Neubarth & Conklin, 2016) is used for mining patterns that contrast between groups of data. The method is designed to discover maximally general distinctive patterns (MGDPs), meaning that for each reported discovered pattern there is no more general pattern that is also distinctive. Each tune family is mined individually for distinctive sequential patterns, using each tune family F as a positive corpus and the rest of the pieces (¬F ) as the anticorpus. In this work a statistical approach is used to measure the distinctiveness of a pattern: it is the probability p of finding at least the observed number of pieces of family F when taking a single random sample of pieces from the entire corpus F ∪ ¬F . A pattern is then considered distinctive if its p-value falls below some specified significance level α (see Conklin, 2013, for details). The MGDP set may contain overlapping patterns, so for the tune family mining task this set is further reduced by a greedy pruning strategy. Proceeding from the best (lowest p-value) pattern, a pattern is placed in the final set if it does not overlap, in any piece, with any pattern already in the final set. Thus none of the patterns in the final set will overlap in any piece with any other pattern.