Classifying Non-speech Vocals: Deep vs Signal Processing Representations

Fatemeh Pishdadian, Prem Seetharaman, Bongjun Kim, Bryan Pardo · 2019

Deep-learning-based audio processing algorithms have become very popular over the past decade.Due to promising results reported for deep-learning-based methods on many tasks, some now argue that signal processing audio representations (e.g.magnitude spectrograms) should be entirely discarded, in favor of learning representations from data using deep networks.In this paper, we compare the effectiveness of representations output by state-of-the-art deep nets trained for task-specific problems, to off-the-shelf signal processing representations applied to those same tasks.We address two tasks: query by vocal imitation and singing technique classification.For query by vocal imitation, experimental results showed deep representations were dominated by signal-processing representations.For singing technique classification, neither approach was clearly dominant.These results indicate it would be premature to abandon traditional signal processing in favor of exclusively using deep networks.

Read the paper · More papers on PaperTik