Solo Voice Detection Via Optimal Cancellation
Christine Smit, Daniel P. W. Ellis · 2007
Automatically identifying sections of solo voices or instruments within a large corpus of music recordings would be useful, for example, to construct a library of isolated instruments to train signal models. We consider several ways to identify these sections, including a baseline classifier trained on conventional speech features. Our best results, achieving frame level precision and recall of around 70%, come from an approach that attempts to track the local periodicity of an assumed solo musical voice, then classifies the segment as a genuine solo or not on the basis of what proportion of the energy can be canceled by a comb filter constructed to remove just that periodicity.