Methods for singing voice identification using energy coefficients as features

Annamari Mesaros, Simina Moldovan · 2006

This paper describes two energy representations of the voice signal and tests their efficiency in singing voice identification. The first set of energy features consists in the Mel-scale energies of 14 frequency bands, covering the whole frequency spectrum of the signal. The second energy representation is obtained by wavelet decomposition of the voice signal. The wavelet and scaling filters for the decomposition are derived from fractional B-spline functions. The wavelet decomposition is done hierarchically, into 14 bands, with octave-band filters, taking into account the specific frequencies of the formants. Both energy representations are tested for singing voice identification on the training set and on unknown data

Read the paper · More papers on PaperTik